Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,247

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5,247 results for “cataloguing”

Learn how ShareScore rates datasets ↗
zenodo56/100

Nabro volcano event catalogue from Lapins et al., 2021, JGR Solid Earth

<p>Catalogue of seismic events&nbsp;from Nabro volcano (Sep 2011 - Oct 2012). Data format is a&nbsp;csv file.</p> <p>Events were detected&nbsp;by U-GPD phase arrival picking model. See following paper for details on event detection and location procedure:&nbsp;<em>A Little Data Goes A Long Way Way:&nbsp;Automating Seismic Phase Arrival Picking at Nabro Volcano With Transfer Learning</em> by Lapins et al., 2021, <a href="https://doi.org/10.1029/2021JB021910">https://doi.org/10.1029/2021JB021910</a>).</p> <p>Original seismic waveforms are from the Nabro Urgency Array (Hammond et&nbsp;al., 2011;&nbsp;<a href="https://doi.org/10.7914/SN/4H_2011">https://doi.org/10.7914/SN/4H_2011</a>), which is publicly available through IRIS Data Services (<a href="http://service.iris.edu/fdsnws/dataselect/1/">http://service.iris.edu/fdsnws/dataselect/1/</a>). See Hammond et&nbsp;al.&nbsp;(<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2021JB021910#jgrb55017-bib-0025">2011</a>) for further details on waveform data access and availability.</p> <p>Full code to reproduce our U-GPD transfer learning model, perform model training, run the U-GPD model over continuous sections of data and use model picks to locate events in NonLinLoc (Lomax et&nbsp;al.,&nbsp;<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2021JB021910#jgrb55017-bib-0044">2000</a>) are available at&nbsp;<a href="https://github.com/sachalapins/U-GPD">https://github.com/sachalapins/U-GPD</a>, with the release (v1.0.0) associated with this study also archived and available through Zenodo (Lapins,&nbsp;<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2021JB021910#jgrb55017-bib-0036">2021</a>;&nbsp;<a href="https://doi.org/10.5281/zenodo.4558121">https://doi.org/10.5281/zenodo.4558121</a>).</p> <p>&nbsp;</p> <p>Dataset column key:</p> <p>time = Origin time of seismic event (UTC)</p> <p>lat = Hypocentre latitude&nbsp;in decimal degrees</p> <p>lon = Hypocentre longitude in decimal degrees</p> <p>depth = Hypocentre depth in km</p> <p>rms = RMS error for phase arrival picks and hypocentre (sec)</p> <p>erh = Estimate of horizontal Gaussian error (km)</p> <p>erz = Estimate of vertical Gaussian error (km)</p> <p>azgap = Azimuthal gap (maximum angle separating two adjacent seismic stations, measured from earthquake&nbsp;epicentre)</p> <p>cluster = HDBSCAN cluster number (see Chapter 6 of&nbsp;Lapins, 2021 doctoral thesis:&nbsp;<em>Detecting and characterising seismicity associated with volcanic and magmatic processes through deep learning and the continuous wavelet transform</em>. Persistent URL: <a href="https://hdl.handle.net/1983/ea90148c-a1b2-47ae-afad-5dd0a8b5ebbd">https://hdl.handle.net/1983/ea90148c-a1b2-47ae-afad-5dd0a8b5ebbd</a>)</p> <p>nab*_p_time = P-wave arrival time for station NAB* (UTC)</p> <p>nab*_p_prob = Maximum detection&nbsp;&#39;probability&#39; around P-wave phase arrival from U-GPD model (between 0 and 1)</p> <p>nab*_s_time = S-wave arrival time for station NAB* (UTC)</p> <p>nab*_s_prob = Maximum detection&nbsp;&#39;probability&#39; around S-wave phase arrival from U-GPD model&nbsp;(between 0 and 1)</p> <p>&nbsp;</p> <p>Station csv column key:</p> <p>Network = Seismic network name</p> <p>Station = Seismic station name</p> <p>Latitude = Latitude in decimal degrees</p> <p>Longitude = Longitude in decimal degrees</p> <p>Elevation_asl_km = Station elevation in km above sea level</p>

opencc-by-4.0Dec 2022View details →
zenodo52/100

Updated DEVOTES indicator catalogue of MSFD indicator systems targeting descriptors D1, D2, D4, and D6

<p>This is version 8 of the Catalogue of Indicators of the FP7 project DEVOTES (grant number 308392) that aims at supporting the implementation of the EU MSFD. This catalogue of indicators is an inventory of existing methods. The metadata have been updated and extended since deiverable D3-1 of the DEVOTES project. You can learn more about the DEVOTES project at: http://www.devotes-project.eu</p> <p>All data are provided without a guarantee of correctness or completeness. We are aware of some errors in the database content and are continuously working on correcting these and on supplementing the content with new metadata.</p> <p>The data can be best viewed with the free DEVOTool software, available at http://www.devotes-project.eu/devotool. By using the DEVOTool software, you accept the license conditions as outlined in section 6 of the software manual distributed together with DEVOTool. </p>

opencc-by-4.0May 2017View details →
zenodo48/100

Catalogue of Diversity of Social Innovation

<p>The &lsquo;Catalogue of Social Innovation Diversity in Rural Areas&rsquo; is the consolidated version of the research database of examples of social innovation in marginalised rural areas developed by the project SIMRA. The file contains a spreadsheet document that includes descriptive information of all the examples reviewed and recorded at some stage in the research database.</p> <p>The catalogue includes basic information for identifying and describing the examples and the characteristics of the social innovation. The total number of examples in the catalogue is 401. Of this number, 243 examples were positively validated using the SIMRA definition of social innovation. The information included in the catalogue for these examples is sufficient to meet the criteria of social innovation as defined by SIMRA. The remaining examples either demonstrate elements of social innovation without meeting all criteria, or include insufficient information to allow a positive validation. Examples can be filtered according to spatial scale, country, sector, topic, form and SIMRA validation.</p>

opencc-by-4.0Mar 2020View details →
zenodo48/100

Catalogues of Semantic Artefacts - Maturity Dimensions and Features

<p>This dataset contains, in three different formats (XSLX, CSV, and PDF), a description of twelve maturity dimensions identified from the literature that can be used to measure the maturity of the semantic artefacts catalogues (SAC). For each dimension, a number from 2 to 6 features has been added, for a total of 43 features overall.</p>

opencc-zeroMay 2023View details →
zenodo48/100

A living catalogue of artificial intelligence datasets and benchmarks for medical decision making

<p>We provide&nbsp;a comprehensive curated catalogue of&nbsp;<strong>artificial intelligence datasets</strong> and <strong>benchmarks for medical decision making</strong>. At the time of first release (April 2021), the dataset contains more than 400&nbsp;biomedical and clinical datasets&nbsp;of which 252 are publicly available or available upon request.</p> <p>The dataset was compiled based on a systematic literature review covering both biomedical and computer science literature and&nbsp;grey literature data sources. All datasets were manually systematized and annotated for meta-information, such as:</p> <ul> <li>Availability and licensing information</li> <li>Type of source data</li> <li>Links to source publications, main references or dataset repositories</li> </ul> <p>Benchmark dataset were additionally annotated for the following information:</p> <ul> <li>Associated task</li> <li>Performance metrics commonly used for evaluation</li> <li>Clinical relevance</li> <li>The availability of data splits</li> </ul> <p>In addition to the versioned TSV file on Zenodo, the dataset can also be explored live via&nbsp;<a href="https://docs.google.com/spreadsheets/d/1QjUxxnZ3tuyW5dj6nkt_o5yJcWUZec4ttfJxO8Zlty4/edit?usp=sharing">this Google Spreadsheet</a>.&nbsp;The dataset is intended as a living, extendable resource. Edit suggestions and additions are encouraged and can be submitted via the comment function of the Google sheet.</p> <p>&nbsp;</p> <p><strong>File descriptions</strong></p> <p><em>annotated-datasets.tsv</em> -- contains the annotated datasets</p> <p><em>arXiv-literature-export.tsv</em> -- contains the original literature record export from arXiv</p> <p><em>pubmed-literature-export.tsv</em> -- contains the original literature record export from PubMed</p> <p><em>README.md</em> -- contains a detailed description of all annotation fields</p>

opencc-by-sa-4.0Apr 2021View details →
zenodo48/100

Aggregated Gut Viral Catalogue (AVrC)

<p>Despite the importance of the gut virome in human health and disease, identifying viral sequences from metagenomic datasets remains computationally challenging. Up to 99% of viral reads lack significant alignments to known viral genomes due to underrepresentation in reference databases. Recent machine learning tools can detect novel viral sequences based on features like k-mer composition or genomic signatures, but are limited to classifying assembled contigs into simplistic viral/non-viral categories. Several large-scale efforts have mined human gut metagenomes to establish viral catalogues, including the Gut Virome Database (33,242 viral OTUs), Cenote Human Virome Database (45,033 OTUs), and Gut Phage Database (142,809 OTUs). However, these catalogues have not been consistently compared for quality, diversity, and completeness. There is an unmet need to harmonize available gut viral sequences into a unified resource for comparing novel viruses against previous efforts. The Aggregated Gut Viral Catalog (AVrC) addresses this gap by harmonizing and aggregating previous mining efforts into a comprehensive resource to allow for the exploration of the Human gut viral diversity and the easier comparison of newly discovered viral sequences.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

SLAMS photometric catalogue

<p>Photometric catalogues from the SLAMS survey, described in &#39;Discovery of a thin stellar stream in the SLAMS survey&#39; (arxiv.org/abs/1711.09103). Columns and units:<br> 1) right ascension J2000 (degrees)<br> 2) declination J2000 (degrees)<br> 3) r-band magnitude (mag)<br> 4) r-band magnitude&nbsp;error&nbsp;(mag)<br> 5) r-band&nbsp;spread_model star/galaxy classifier<br> 6) r-band spread_model error<br> 7) g-band magnitude (mag)<br> 8) g-band magnitude&nbsp;error&nbsp;(mag)<br> 9) g-band&nbsp;spread_model star/galaxy classifier<br> 10) g-band spread_model error<br> 11) E(B-V) reddening (mag)</p>

opencc-by-4.0Aug 2018View details →
zenodo48/100

AMnrGC - Amazon river non-reduntant microbial genes catalogue

<p>&nbsp;</p> <p><strong>AMnrGC : Amazon river basin non-redundant microbial gene catalogue</strong></p> <p>&nbsp;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; RELEASE 2018/01<br> &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; --------------------------------------</p> <p>&nbsp;</p> <p>1. INTRODUCTION</p> <p>&nbsp;&nbsp; AMnrGC is a collection of genes and proteins which were constructed<br> &nbsp;&nbsp; by use of Amazon river basin openly available metagenomes from<br> &nbsp;&nbsp; sequencing projects (SRP044326, PRJEB25171 and SRP039390). Briefly,<br> &nbsp;&nbsp; metagenomes were coassembled by groups made up their geographical<br> &nbsp;&nbsp; location with Megahit v.1.0 and the contigs were used to gene predictions<br> &nbsp;&nbsp; by Prodigal v.2.6.3. Genes sequences were length filtered (&gt; 150 bp) and<br> &nbsp;&nbsp; clustered by CD-HIT-EST (version 4.6) at 95% of nucleotide identity and<br> &nbsp;&nbsp; 90% of overlap of the shorter gene. Theorical protein products were annotated<br> &nbsp;&nbsp; by the most completes databases up to date and their complete information<br> &nbsp;&nbsp; is available here.</p> <p>&nbsp;</p> <p>2. LOCATION</p> <p>&nbsp;&nbsp; AMnrGC versions will be available on the web only under the current ZENODO<br> &nbsp;&nbsp; repository: 10.5281/zenodo.1484504</p> <p>&nbsp;</p> <p>3. FORMAT</p> <p>&nbsp;&nbsp; Gene entries were named as &quot;&gt;AM_AGSSY_XXX&quot; where XXX represents an unique numerical<br> &nbsp;&nbsp; identifier. Genes were deposited in their coding phase, because of this, all of them<br> &nbsp;&nbsp; can be used to generate the protein sequences by transeq function at ORF+1.<br> &nbsp;&nbsp; The protein entries correspond to genes artifical translation used in the annotations,<br> &nbsp;&nbsp; and also available, codified in the same way, but containing the indication &quot;_1&quot; in<br> &nbsp;&nbsp; the end of the header. Example:</p> <p>&nbsp;&nbsp; &nbsp;Gene:<br> &nbsp;&nbsp; &nbsp;&gt;AM_AGSSY_151515</p> <p>&nbsp;&nbsp; &nbsp;Protein:<br> &nbsp;&nbsp; &nbsp;&gt;AM_AGSSY_151515_1</p> <p>&nbsp;&nbsp; Annotations were provided as separate tables for each database used to annotate the<br> &nbsp;&nbsp; sequences. The header of these tables indicates the meaning of each value.</p> <p>&nbsp;</p> <p>2. FUTURE FORMAT CHANGES</p> <p>&nbsp;&nbsp; No major changes are expected for the main general format of the database.<br> &nbsp;&nbsp; New versions should include updated versions of annotations or even additional sequences,<br> &nbsp;&nbsp; numbered as subsequent entries.</p> <p>&nbsp;</p> <p>3. ACKNOWLEDGEMENTS<br> &nbsp; &nbsp;<br> &nbsp;&nbsp; This work is a joint effort of Laboratory of molecular biology from Federal<br> &nbsp;&nbsp; University of S&atilde;o Carlos, S&atilde;o Paulo, Brazil (LBM/UFSCAR) and Protists group<br> &nbsp;&nbsp; of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> &nbsp;&nbsp; Conselho Nacional de Desenvolvimento Cient&iacute;fico e Tecnol&oacute;gico (CNPq), as well as, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; the spanish funding organ Consejo Superior de Investigaciones Cient&iacute;ficas (CSIC).</p> <p>&nbsp;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; This study was financed in part by the Coordena&ccedil;&atilde;o de Aperfei&ccedil;oamento de Pessoal &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; de N&iacute;vel Superior - Brasil (CAPES) - Finance Code 001.</p> <p>&nbsp;</p> <p>4. THE AMnrGC TEAM</p> <p>&nbsp;&nbsp; AMnrGC is maintained by a group of researchers. You can contact<br> &nbsp;&nbsp; the AMnrGC consortium.<br> &nbsp; &nbsp;<br> &nbsp;&nbsp; Current curators:</p> <p>&nbsp;&nbsp; - C&eacute;lio Dias Santos J&uacute;nior (celio.diasjunior@gmail.com)<br> &nbsp;&nbsp; - Flavio Henrique-Silva (dfhs@ufscar.br)<br> &nbsp;&nbsp; - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> &nbsp;</p> <p>5. COPYRIGHT NOTICE</p> <p>&nbsp;&nbsp; AMnrGC - Amazon river basin non-redundant microbial gene catalogue<br> &nbsp;&nbsp; Copyright (C) 2018 The AMnrGC consortium.</p> <p>&nbsp;&nbsp; This database is provided &ldquo;as is&rdquo; and without any warranty of any kind,<br> &nbsp;&nbsp; of openly available for non-commerical purposes. You can redistribute and/or modify it<br> &nbsp;&nbsp; as you wish, under the terms of the ODBL 1.0 license:</p> <p>&nbsp;&nbsp; &nbsp;https://opendatacommons.org/licenses/odbl/1.0/</p> <p>&nbsp;&nbsp; For commercial purposes, please contact us. &nbsp;</p> <p>___________________</p> <p>The AMnrGC Consortium<br> 2018</p>

opencc-by-4.0Nov 2018View details →
zenodo48/100

Identifying Coronal Mass Ejection Active Region Sources: An automated approach - Catalogue results

<p>Catalogue of Coronal Mass Ejection (CME) active region sources. Includes a database version and a simplified .csv version. For full details, refer to the source code at <a href="https://github.com/JulioHC00/cmesrc">https://github.com/JulioHC00/cmesrc</a>. We include a README file for each describing each column.</p> <p>We also include the raw data used to generate the catalogue so that results may be reproduced following the steps detailed in <a href="https://github.com/JulioHC00/cmesrc">https://github.com/JulioHC00/cmesrc</a>. This is a collection of data from other works and we provide it only to allow the results to be reproduced</p> <p>Below, we detail the data sources for the raw_data folders</p> <p>==============================<br><strong>RAW DATA SOURCES</strong><br>==============================</p> <p><strong>DIMMINGS FOLDER</strong></p> <p>Data is from Solar Demon, .csv was provided by Emil Kraaikamp through private communication.</p> <blockquote> <p>Solar Demon &ndash; an approach to detecting flares, dimmings, and EUV waves on SDO/AIA images<br>Emil Kraaikamp, Cis Verbeeck<br>J. Space Weather Space Clim. 5 A18 (2015)<br>DOI: 10.1051/swsc/2015019</p> </blockquote> <p><strong>HARPNUM_TO_NOAA FOLDER</strong></p> <p>Obtained from http://jsoc.stanford.edu/doc/data/hmi/harpnum_to_noaa/all_harps_with_noaa_ars.txt</p> <p><strong>LASCO FOLDER</strong></p> <p>This CME catalog is generated and maintained at the CDAW Data Center by NASA and The Catholic University of America in cooperation with the Naval Research Laboratory. SOHO is a project of international cooperation between ESA and NASA.</p> <p>Downloaded from https://cdaw.gsfc.nasa.gov/CME_list/</p> <p><strong>MVTS FOLDER</strong></p> <p>Data from</p> <blockquote> <p>Angryk, R.A., Martens, P.C., Aydin, B. et al. Multivariate time series dataset for space weather data analytics. Sci Data 7, 227 (2020). https://doi.org/10.1038/s41597-020-0548-x</p> </blockquote> <p>Available at the Harvard Dataverse</p> <blockquote> <p>Angryk, Rafal; Martens, Petrus; Aydin, Berkay; Kempton, Dustin; Mahajan, Sushant; Basodi, Sunitha; Ahmadzadeh, Azim; Xumin Cai; Filali Boubrahimi, Soukaina; Hamdi, Shah Muhammad; Schuh, Micheal; Georgoulis, Manolis, 2020, "SWAN-SF", https://doi.org/10.7910/DVN/EBCFKM, Harvard Dataverse, V1</p> </blockquote> <p>The DT_SWAN folder contains the same data but with extra columns obtained directly from the Joint Science Operations Center (JSOC) through the python package drms.</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

UniCat library catalogue dataset

<p>This dataset was created in the context of a project about libcitations or library catalogue analysis of book publications by Flemish Social Sciences and Humanities researchers.&nbsp;</p> <p>The dataset was constructed on the basis of&nbsp;the openly available API of the Belgian UniCat library catalogue (<a href="https://www.unicat.be/uniCat?func=search&amp;uiLanguage=en">UniCat-Search</a>). The library catalogue was searched in September 2021 by a matching of the catalogue against the ISBNs present in the VABB-SHW database (see&nbsp;Aspeslagh, Peter, Guns, Raf, &amp; Engels, Tim C.E. (2021). VABB-SHW: Dataset of Flemish Academic Bibliography for the Social Sciences and Humanities (edition 11) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5795899).</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

Piano sonatas and Beethoven works in Hans Georg Nägelis catalogues

<p>This dataset (in two csv files) collects the references to piano sonatas, and to works by Ludwig van Beethoven, in the catalogues published by Hans Georg N&auml;geli as a bookseller in Zurich between 1792 and 1805.</p> <p>The data was used by the author in his paper &quot;Hans Georg N&auml;geli as Publisher and Bookseller of Piano Music&quot;, presented at the conference&nbsp;<a href="https://www.hkb-interpretation.ch/beethoven2020">Beethoven and the Piano</a>,&nbsp;4&ndash;7&nbsp;November 2020.</p> <p>The printed version of the article is at present in preparation.</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Sentinel-2 Cloud Mask Catalogue

<p><strong>Overview</strong></p> <p>This dataset comprises cloud masks for 513 1022-by-1022 pixel subscenes, at 20m resolution, sampled random from the 2018 Level-1C Sentinel-2 archive. The design of this dataset follows from some observations about cloud masking: (i) performance over an entire product is highly correlated, thus subscenes provide more value per-pixel than full scenes, (ii) current cloud masking datasets often focus on specific regions, or hand-select the products used, which introduces a bias into the dataset that is not representative of the real-world data, (iii) cloud mask performance appears to be highly correlated to surface type and cloud structure, so testing should include analysis of failure modes in relation to these variables.</p> <p>The data was annotated semi-automatically, using the <a href="https://github.com/ESA-PhiLab/iris">IRIS toolkit</a>, which allows users to dynamically train a Random Forest (implemented using <a href="https://github.com/microsoft/LightGBM">LightGBM</a>), speeding up annotations by iteratively improving it&#39;s predictions, but preserving the annotator&#39;s ability to make final manual changes when needed. This hybrid approach allowed us to process many more masks than would have been possible manually, which we felt was vital in creating a large enough dataset to approximate the statistics of the whole Sentinel-2 archive.</p> <p>In addition to the pixel-wise, 3 class (CLEAR, CLOUD, CLOUD_SHADOW) segmentation masks, we also provide users with binary<br> classification &quot;tags&quot; for each subscene that can be used in testing to determine performance in specific circumstances. These include:</p> <ul> <li><strong>SURFACE TYPE</strong>: <em>11 categories</em></li> <li><strong>CLOUD TYPE</strong>: <em>7 categories</em></li> <li><strong>CLOUD HEIGHT</strong>: <em>low, high</em></li> <li><strong>CLOUD THICKNESS</strong>: <em>thin, thick</em></li> <li><strong>CLOUD EXTENT</strong>: <em>isolated, extended</em></li> </ul> <p>&nbsp;</p> <p>Wherever practical, cloud shadows were also annotated, however this was sometimes not possible due to high-relief terrain, or large ambiguities. In total, 424 were marked with shadows (if present), and 89 have shadows that were not annotatable due to very ambiguous shadow boundaries, or terrain that cast significant shadows. If users wish to train an algorithm specifically for cloud shadow masks, we advise them to remove those 89 images for which shadow was not possible, however, bear in mind that this will systematically reduce the difficulty of the shadow class compared to real-world use, as these contain the most difficult shadow examples.</p> <p>In addition to the 20m sampled subscenes and masks, we also provide users with shapefiles that define the boundary of the mask on the original Sentinel-2 scene. If users wish to retrieve the L1C bands at their original resolutions, they can use these to do so.</p> <p>Please see the README for further details on the dataset structure&nbsp;and more.</p> <p>&nbsp;</p> <p><strong>Contributions &amp; Acknowledgements</strong></p> <p>The data were collected, annotated, checked, formatted and published by Alistair Francis and John Mrziglod.</p> <p>Support and advice was provided by Prof. Jan-Peter Muller and Dr. Panagiotis Sidiropoulos, for which we are grateful.</p> <p>We would like to extend our thanks to Dr. Pierre-Philippe Mathieu and the rest of the team at <em>ESA PhiLab</em>, who provided the environment in which this project was conceived, and continued to give technical support throughout.</p> <p>Finally, we thank the <em>ESA Network of Resources</em> for sponsoring this project by providing ICT resources.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

SCORPIO ASKAP15 compact radio source catalogue

<p>This dataset provides the compact radio source catalogue data&nbsp;extracted from ASKAP observations (15 antennae) of the SCORPIO field at 912 MHz, carried out in the context of ASKAP EMU Early Science phase. The reference scientific publication is S.Riggi et al., to appear on MNRAS.&nbsp;</p> <p>The dataset includes:</p> <p>-&nbsp; Source catalogue in tabular format: Two ascii/FITS table files with a series of summary parameters for each catalogued source islands and fitted components, respectively. Table format (number of data columns and column description) is detailed in the CAESAR source finder online documentation at <a href="http://caesar-doc.readthedocs.io">https://caesar-doc.readthedocs.io</a>. Additionally, we provide an added-value source component catalogue table (ascii and FITS formats) with extra-information (corrected fluxes, radio/infrared cross-match info, spectral indices, etc.). Its format is described in the reference publication;</p> <p>-&nbsp;Source catalogue in ROOT format:&nbsp;A ROOT&nbsp;file storing the list of catalogued sources and relative components as a CAESAR <em>Source</em>&nbsp;C++ object. For each source the summary parameters plus detailed information at pixel level are available. The detailed format is described in the CAESAR API documentation at <a href="http://caesar-doc.readthedocs.io">https://caesar-doc.readthedocs.io</a>.</p> <p>-&nbsp;Source list in region format: Two DS9 region files with the list of catalogued source islands and fitted components, respectively reported as labelled polygons or ellipses.</p> <p>- Background maps in FITS format: Two FITS files with background and noise maps obtained in the source finding process.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

3D CMT catalogue of moderate size offshore earthquakes along the Nankai Trough

<p>3D CMT inversion solutions of moderate-size earthquakes along the Nankai Trough, <strong>version 3.1.&nbsp;</strong>&nbsp;</p> <ul> <li>Analyzed periods:&nbsp;<strong>January 2003&nbsp;to December&nbsp;2020</strong> <ul> <li>The catalog version 3, containing CMT solutions from January 2003 to April 2020.</li> <li>The catalog version 2.2, which is containing CMT solutions from April 2004 to August 2019, has been published in GJI (Takemura, Okuwaki et al. 2020&nbsp;<a href="https://doi.org/10.1093/gji/ggaa238">doi:10.1093/gji/ggaa238</a>).</li> </ul> </li> <li>The method is&nbsp;described in Takemura, Okuwaki, et al., 2020, GJI, <a href="https://doi.org/10.1093/gji/ggaa238">doi:10.1093/gji/ggaa238</a>&nbsp;<a href="https://doi.org/10.31223/osf.io/nbd79">the submitted preprint</a>.&nbsp;</li> </ul> <p>If you use this version, you should cite the appropriate DOI and Takemura, Okuwaki, et al. 2020 GJI.</p> <p><strong>Included files</strong></p> <ul> <li>YYYYMMDDHHMM_25-100s__CMT.dat<br> CMT solutions at all selected source grids for an earthquake that occurred at HH:MM on DDth MM YYYY (JST). Latitude, longitude, depth, VR [%], M<sub>rr</sub>, M<sub>tt</sub>, M<sub>ff</sub>, M<sub>rt</sub>, M<sub>rf</sub>, M<sub>tf</sub>, exponent (dyne-cm), Mo [Nm], strike1, dip1, rake1, strike2, dip2, rake2, Mw, index of source grid (internal parameter), and centroid time are listed.&nbsp;</li> <li>YYYYMMDDHHMM_25-100s__CMTparam.dat<br> Input directory (internal parameter), Green&#39;s function directory (internal parameter), the number of source grids, the number of used stations, station names used in CMT inversion, frequency range, initial epicenter and distance range are listed.</li> <li>3DCMTcatalog_v3.csv<br> CSV format file of the 3D CMT catalog for earthquakes with Mw of 4.3-6.5</li> <li>catalog3DCMT_Takemura2019_Mw7.2_7.5SEKii.csv<br> CSV format file of the 3D CMT catalog for the Mw 7.2 and 7.5 southeast off the Kii Peninsula earthquake occurred on 19:07 and 23:57 5th September 2004 (JST), respectively.</li> </ul>

opencc-by-4.0Oct 2019View details →
zenodo44/100

Example Instantiations of the e-Infrastructure Catalogue of Services

<p>This document contains example instantiations of the e-Infrastructure Catalogue of Services. This supports a published framework for creating a Catalogue of Services (CoS), primarily intended for e-Infrastructure services, which describes services at a high level and makes them discoverable.</p>

opencc-by-4.0Nov 2016View details →
zenodo44/100

Model catalogues and histograms of KSVZ axion models with multiple heavy quarks

<p>This record contains the files pertaining to the results mentioned in the linked work Plakkot &amp; Hoof, <em>&ldquo;Anomaly Ratio Distributions of KSVZ Axion Models with Multiple Heavy Quarks.&rdquo;</em> The contents are:</p> <ul> <li>4 histogram files</li> <li>2 compressed <code>tar.gz</code> files containing detailed catalogues,</li> <li>2 Python scripts to extract information from the files (requires Python 3 and the <code>numpy</code>, <code>matplotlib</code> and <code>h5py</code> packages)</li> </ul> <p><strong>Histogram files</strong></p> <p>The columns of the histogram files (listed below) correspond to the numerator of <em>E</em>, denominator of <em>E</em>, numerator of <em>N</em>, denominator of <em>N</em>, and the number of models (frequency) with that <em>E/N</em> ratio. The file <code>histogram_complete_NQ_1_to_9.txt</code> contains additional columns showing the number of models per entry for each <em>N<sub>Q</sub></em>.</p> <ol> <li><code>histogram_all_LP_allowed_models.txt</code> for all LP-allowed models</li> <li><code>histogram_additive_LP_allowed_models.txt</code> for LP-allowed additive models</li> <li><code>histogram_same_reps_LP_allowed_models.txt</code> for LP-allowed additive models where all new quarks live on the same representation</li> <li><code>histogram_complete_NQ_1_to_9.txt</code> for all possible models with <em>N<sub>Q</sub></em> &le; 9, regardless of the LP criterion</li> </ol> <p><strong>Catalogues</strong></p> <p>The catalogues are contained in the two compressed <code>HDF5</code> files listed below. The catalogues contain groups for different <em>N<sub>Q</sub></em>, each with subgroups for numerators and denominators of E and N, model representation (as a list). The file <code>catalogue_additive_models_NQ_1_to_9.tar.gz</code> contains additionally the energy scale at which the first LP appears. The integers in the model lists represent the quark representations (integer <em>m</em> for the representation <em>r<sub>m</sub></em> in the text), and the negative signs indicate opposite PQ charge.</p> <ol> <li><code>catalogue_additive_models_NQ_1_to_9.tar.gz</code> contains the catalogue for additive models with <em>N<sub>Q</sub></em> &le; 9 (9 groups, 6 subgroups)</li> <li><code>catalogue_all_LP_allowed_models.tar.gz</code> Contains catalogues for all LP-allowed models (28 groups, 5 subgroups)</li> </ol> <p><strong>Scripts</strong></p> <p>The two scripts are:</p> <ol> <li><code>print_histogram_info.py</code> is a sample script to extract information from the histogram files. The default choice for <em>N<sub>Q</sub></em> can be adjusted by e.g. invoking <code>python print_histogram_info.py 5</code> for <em>N<sub>Q</sub></em> = 5</li> <li><code>read_catalogue.py</code> is a sample script to extract data from the catalogue files (which need to be unpacked first). The default settings can be overwritten by e.g. invoking <code>python read_catalogue.py 5 file.h5</code> for <em>N<sub>Q</sub></em> = 5 and the catalogues contained in <code>file.h5</code></li> </ol>

opencc-by-4.0Jul 2021View details →
zenodo44/100

e-ITALICA (enhanced ITAlian rainfall-induced LandslIdes CAtalogue)

<p>The enhanced ITAlian rainfall-induced LandslIdes CAtalogue (e-ITALICA) currently lists 6312 records with information on rainfall-induced landslides that occurred over the Italian territory between January 1996 and December 2021. Information on rainfall-induced landslides has a high accuracy on their spatial and temporal location. e-ITALICA includes the triggering rainfall conditions associated with the landslides and the coordinates of the representative rain gauges. Moreover, details on elevation, slope, and land cover are also included.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

A Catalogue of All Known Mass-Transferring Ultracompact Binary Systems

<p>Introduction: I present a catalogue that collects all known mass-transferring ultracompact binary systems in one place. This includes AM CVn-type binaries, helium-enriched CVs and a variety of other objects.</p> <p>The goal of this catalogue is to prevent the duplicated effort of multiple researchers searching for known systems through the literature. Where applicable, the catalogue includes orbital periods, measured masses, Gaia cross-matches, as well as important references for each system. The catalogue includes 'confirmed' systems (usually means an orbital period measurement and/or a spectrum) and 'candidates' (may be selected based on other properties such as outburst shape).</p> <p>For full details and references see the associated paper (accepted to A&amp;A). Preprint at https://arxiv.org/abs/2505.10535</p> <p>See the catalogue itself in the file 'amcvn_catalogue.fits'.</p> <p>Up to date as of 2025-04-01.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Catalogue of Plants 2021. Royal Botanic Garden Edinburgh (data)

<p>This is a static snapshot of the Royal Botanic Garden Edinburgh (RBGE) Living Collection Catalogue held in the Living Collections Management System (CMS) it is produced as a record of what was being grown at the RBGE on 10/09/2021.&nbsp;</p> <p>The data includes the specialist gardens of Benmore Botanic Garden (BBG), Dawyck Botanic Garden (DBG), Logan Botanic Garden (LBG), Sites of the &nbsp;International Conifer Conservation Programme (ICCP) and the Scottish Native Plants Project.</p> <p>The dynamic (up to date) data available at</p> <p><a href="https://data.rbge.org.uk/search/livingcollection/">https://data.rbge.org.uk/search/livingcollection/</a></p> <p>&nbsp;</p> <p><strong>How to Use the Catalogue Spreadsheet</strong></p> <p><em>By Benedict Lyte</em></p> <p>The catalogue is divided into nine sections:</p> <ol> <li>Bryophytes</li> <li>Fern allies</li> <li>Ferns</li> <li>Gnetophytes</li> <li>Conifers</li> <li>Ginkgophytes</li> <li>Cycads</li> <li>Dicotyledons</li> <li>Monocotyledons</li> </ol> <p>&nbsp;</p> <p>The listing is organised alphabetically by family following APGIV (<a href="http://www.mobot.org/MOBOT/research/APweb">http://www.mobot.org/MOBOT/research/APweb</a>) and then accession number.</p> <p>The taxon data uses World Flora Online (<a href="http://www.worldfloraonline.org">http://www.worldfloraonline.org</a>) for verification.</p> <p>&nbsp;</p> <p><strong>Accession number</strong></p> <p>RBGE accession numbers are eight digits long: the first four digits represent the year of the accession and are followed by the sequential number of that accession in that year.</p> <p>&nbsp;</p> <p><strong>The location where the accession is alive</strong></p> <p>BBG = Benmore Botanic Garden</p> <p>DBG = Dawyck Botanic Garden</p> <p>InvI = Inverleith, under glass</p> <p>InvO = Inverleith, outdoors</p> <p>LBG = Logan Botanic Garden</p> <p>Ext = External site that is part of the International Conifer Conservation Programme or the Scottish Native Plants Project</p> <p>&nbsp;</p> <p><strong>Country of origin</strong></p> <p>Names of the countries follow those given by the International Organization for Standardization standard 3166 (ISO, 2011).</p> <p>&nbsp;</p> <p><strong>Collection info</strong></p> <p>Collection ID is a code allocated by RBGE.</p> <p>Collector is the name(s) of the individual(s) on the collection trip with that ID.</p> <p>Collection number is the number assigned to the accession when it was collected.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

2019 search and interaction log from the data catalogue: Research Data Australia

<p>In order to provide a better support to user&#39;s data discovery activity, we analysed a data search log in order to understand how data seekers interact with a data search system when they search for data.&nbsp; The data search log is from the research data discovery portal: <a href="https://researchdata.edu.au">Research Data Australia (RDA)</a>. RDA&nbsp; is the data discovery service of the Australian Research Data Commons (ARDC). ARDC is supported by the Australian Government through the National Collaborative Research Infrastructure Strategy Program.</p> <p>Please read the research paper &quot;<a href="https://doi.org/10.1108/JD-12-2021-0245">Large-scale Analysis of Query Logs to Profile Users for Dataset Search</a>&quot; for detailed description and analysis of the datasets, and the software &quot;<a href="https://zenodo.org/record/6321621#.Yh79Tt9xUmA">Python code for processing and clustering a data search log</a>&quot; for the data process and analysis.</p> <p>The search log consists of the entire user-front activity log data for the duration of January to December 2019.&nbsp; During this period, the catalogue contained about 150,000 metadata records of datasets.</p> <p>The dataset (2019_search_log_sessioned.txt) was generated from raw log data with following steps:</p> <ul> <li>Remove entries that were likely from machines instead of human users. Those recorded machine activities may result from downstream aggregators who harvested metadata from RDA by directly sending queries to the catalogue URL instead of using the API endpoint.</li> <li>Identify search sessions from a user - a search session includes all activities a user conducts with a search system in order to satisfy a (information/data) search needs.&nbsp; We followed the following steps to identify search sessions. First, we identified a user by IP address, where a unique IP address was considered a single user. We recognise the limitation of this approach, as several users may share the same IP address, however the IP address is the only information available for identifying a user.&nbsp;<br> Past research in log analysis usually apply the following two methods to identify a session: 30 minutes from the same IP address, and/or more than 30 minutes of inactivity between the current activity event and its immediate preceding event. We examined both methods carefully for our log data and concluded that both ended with large unwanted sessions from machine activities. Therefore, we take a brutal approach, by taking only a session from an IP address with a maximum 30 minutes duration.</li> <li>We also removed sessions whose 40% of activities resulted in &rsquo;page not found&rsquo; or whose activities were all about accessing grants. Within a session, we removed &quot;duplicated&quot; activities that were exactly as their precedent activity with less than one second time span (this could have been a result of reloading a page).</li> </ul> <p>The dataset (id_to_title_subject.csv) lists title and subject headings per record id.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record