Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,655

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,655 results for “SubSet”

Learn how ShareScore rates datasets ↗
zenodo44/100

Wikidata Subsetting: Reference-based Subsetting Experiment Datasets

<p>Files in this dataset have been produced during Flexibility experiments of Wikidata subsetting practical tools: Subsetting based on references using WDSub.</p>

opencc-byJun 2023View details →
zenodo44/100

Wikidata Subsetting: Performance and Accuracy Experiment Datasets

<p>Files in this dataset have been produced during Performance and Accuracy experiments of Wikidata subsetting practical tools: WDumper, KGTK, WDSub, WDF.</p>

opencc-byJun 2023View details →
zenodo44/100

zbMATH Open Access Subset

<p>The dataset contains two tables as csv files.</p> <p>1) documents_in_oa_series</p> <p>is a list of zbMath Open documents in serials where the description contains the word &quot;Open Access&quot;. Note that some documents did not appear yet or might have been retracted. Thus when fetching information from the oai-pmh API or the website, be prepared to handle non-existing documents.</p> <p>2) oa_links</p> <p>Lists all links from zbMATH Open documents to fulltext matched by either unpaywall or arxiv.</p> <p>The meaning of the fields is:</p> <ul> <li><strong>zbmath_id</strong> Unique identifier from zbMATH Open. Prefix with <code>https://zbmath.org/</code> to visit additional information on the article. For example, <code>5635019</code> is associated with <a href="https://zbmath.org/5635019">https://zbmath.org/5635019</a></li> <li><strong>link</strong> link to the fulltext. Note that those links are not fully reliable. We estimate a success rate of 90%</li> </ul> <p>Note that there is an overlap between 1 and 2. So some documents are published in OA serials and have arxiv or unpaywall links at the same time.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Subset of 'MLSUM: The Multilingual Summarization Corpus' for constraints annotation experiment

<p><strong>[EN] Subset of &#39;MLSUM: The Multilingual Summarization Corpus&#39; for constraints annotation experiment.</strong></p> <ul> <li><strong>Description</strong>: MLSUM is a dataset of newspappers articles aimed at training summaring model. We use it for a constraints annotation experiment on newspapper titles according to their topic classification.</li> <li><strong>Content</strong>: For constraints annotation experiment based on data similarity, this dataset have been subsetted (randomly pick 75 articles in the following 14 most used topics: &#39;economie&#39;, &#39;politique&#39;, &#39;sport&#39;, &#39;planete&#39; (renamed in &#39;ecologie&#39;), &#39;sciences&#39;, &#39;police-justice&#39;, &#39;disparitions&#39;, &#39;emploi&#39;, &#39;sante&#39;, &#39;musiques&#39;, &#39;arts&#39;, &#39;educations&#39;, &#39;climat&#39; (renamed in &#39;meteo&#39;), &#39;immobilier&#39;) and filtered (keep articles that have an obvious topics regarding their titles, without their bodies). Two reviewers have working on this task in order to limit the subjectivity of the filtering. This subsetted dataset is used (1) to estimate needed time to annotate titles similarity with constraints (MUST-LINK, CANNOT-LINK) and (2) to test interactive clustering methodology (constraints annotation and constrained clustering).</li> <li><strong>Origin</strong>: The dataset is bassed on the original &#39;MLSUM: The Multilingual Summarization Corpus&#39; dataset (https://doi.org/10.48550/arXiv.2004.14900).</li> </ul> <p><br> <strong>[FR] Echantillon de &#39;MLSUM: The Multilingual Summarization Corpus&#39; pour une exp&eacute;rience&nbsp;d&#39;annotation de contraintes.</strong></p> <ul> <li><strong>Description </strong>: MLSUM est un ensemble de donn&eacute;es d&#39;articles de journaux destin&eacute;s &agrave; l&#39;entra&icirc;nement d&#39;un mod&egrave;le de r&eacute;sum&eacute; automatique. Nous l&#39;utilisons pour une exp&eacute;rience d&#39;annotation de contraintes sur des titres de journaux en fonction de leur classification th&eacute;matique.</li> <li><strong>Contenu </strong>: Pour une exp&eacute;rience d&#39;annotation de contraintes bas&eacute;e sur la similarit&eacute; des donn&eacute;es, cet ensemble de donn&eacute;es a &eacute;t&eacute; &eacute;chantillonn&eacute; (s&eacute;lectionner au hasard de 75 articles dans les 14 sujets les plus utilis&eacute;s&nbsp;: &#39;&eacute;conomie&#39;, &#39;politique&#39;, &#39;sport&#39;, &#39;plan&egrave;te&#39; (renomm&eacute; en &laquo; &eacute;cologie &raquo;). ), &#39;sciences&#39;, &#39;police-justice&#39;, &#39;disparitions&#39;, &#39;emploi&#39;, &#39;sante&#39;, &#39;musiques&#39;, &#39;arts&#39;, &#39;&eacute;ducations&#39;, &#39;climat&#39; (renomm&eacute; en &#39;meteo&#39;), &#39;immobilier&#39; ) et filtr&eacute; (conserver les articles qui ont un sujet &eacute;vident par rapport &agrave; leur titre, sans leur corps). Deux relecteurs ont travaill&eacute; sur cette t&acirc;che afin de limiter la subjectivit&eacute; du filtrage. Ce sous-ensemble de donn&eacute;es est utilis&eacute; (1) pour estimer le temps n&eacute;cessaire pour annoter la similarit&eacute; des titres avec des contraintes (MUST-LINK, CANNOT-LINK) et (2) pour tester la m&eacute;thodologie de clustering interactif (annotation de contraintes et clustering contraint).</li> <li><strong>Origine </strong>: L&#39;ensemble de donn&eacute;es est bas&eacute; sur l&#39;ensemble de donn&eacute;es original &#39;MLSUM : The Multilingual Summarization Corpus&#39; (https://doi.org/10.48550/arXiv.2004.1490).</li> </ul>

openmit-licenseOct 2023View details →
zenodo40/100

Subset of nucleosomal DNA sequences from mouse brain nucleus accumbens tissue (GEO dataset GSE54263)

<p>This dataset contains a subset of nucleosomal DNA sequences of +1 nucleosomes from mouse brain nucleus accumbens cells (NAC) used to analyze nucleosome positioning sequence (NPS) patterns in&nbsp;<a href="https://doi.org/10.1371/journal.pcbi.1007365">Pranckeviciene, Erinija and Hosid, Sergey and Liang, Nathan and Ioshikhes, Ilya (2020). Nucleosome positioning sequence patterns as packing or regulatory. In PLoS computational biology, 16 (1), pp. e1007365.</a></p> <ul> <li>controlm.fa.gz contains sequences of <strong>control</strong> mice (GSE54263 subset Con_H3 GSM1311267)</li> <li>&nbsp;resilientm.fa.gz contains sequences of mice <strong>resilient to social stress</strong> (GSE54263 subset Res_H3 GSM1311268)</li> <li>&nbsp;susceptiblem.fa.gz contains sequences of<strong> </strong>mice <strong>susceptible to social stress</strong> (GSE54263 subset Sus_H3 GSM1311269)</li> </ul> <p>This dataset originates from the GEO accession GSE54263 data from <a href="https://www.nature.com/articles/nm.3939">Sun H, Damez-Werno DM, Scobie KN, Shao NY et al. ACF chromatin-remodeling complex mediates stress-induced depressive-like behavior. <em>Nat Med</em> 2015 Oct;21(10):1146-53.</a></p>

opencc-by-4.0May 2020View details →
zenodo40/100

ESSENTIA analysis of audio snippets from the Million Song Dataset Taste Profile subset

<p>This upload includes the ESSENTIA analysis output of (a subset of) song snippets from the Million Song Dataset, namely those included in the Taste Profile subset. The audio snippets were collected from 7digital.com and were subsequently analyzed with ESSENTIA 2.1-beta3. Pre-trained SVM models provided by the ESSENTIA authors on their website were applied.</p> <p>The file <strong>msd_song_jsons.rar </strong>contains the ESSENTIA analysis output after applying the SVM models for highlevel feature extraction. Please note that these are 204317 files.</p> <p>The file <strong>msd_played_songs_essentia.csv.gz </strong>contains all one-dimensional real-valued fields of the jsons merged into one csv file with 204317 rows.</p> <p>The full procedure and subsequent analysis is described in</p> <p>Fricke, K. R., Greenberg, D. M., Rentfrow, P. J., &amp; Herzberg, P. Y. (2019). Measuring musical preferences from listening behavior: Data from one million people and 200,000 songs. <em>Psychology of Music</em>, 0305735619868280.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Dataset for "Intraspecies diversity reveals a subset of highly variable plant immune receptors and predicts their binding sites"

<p>Datasets for preprint (https://doi.org/10.1101/2020.07.10.190785) entitled:</p> <p>&quot;Intraspecies diversity reveals a subset of highly variable plant immune receptors and predicts their binding sites&quot;</p> <p>Contains:</p> <p>- All data&nbsp;files for scripts quoted in the preprint and&nbsp;deposited at&nbsp;https://github.com/krasileva-group/hvNLR</p> <p>- Clade Membership Tables</p> <p>- Clade Alignment Files</p> <p>- Clade Trees</p> <p>- Excel&nbsp;files for Figure S1, Figure S2, and Table 1</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Wikidata Subsets and Specification Files Created by WDumper

<p>WDumper is a third-party tool enables users to craete custom dump of Wikidata. Here we create some topical subset of Wikidata by WDumper. There are 4 usecases:</p> <ol> <li>politicians: People with occupation of politicians in Wikidata</li> <li>militPoliticians: People with occupation of politicians which are military and also have a military rank of General in Wikidata</li> <li>ukUniversities: All United Kingdom universities in Wikidata</li> <li>geneWiki: A subset from <a href="https://elifesciences.org/articles/52614/figures#fig1">this class-diagram</a></li> </ol> <p>The subsets are in .nt.gz format. For each usecase there are two types of subsets in terms of content, one with References and Qualifiers that has a &quot;withRQFS&quot; in the name, and one without this feature. Also for each use case, there are two extracted subsets one from 27 April 2015 and another from 13 November 2020. The corresponding JSON file for each use case is the WDumper specification file.</p>

opencc-by-4.0Feb 2021View details →
zenodo40/100

A Subset of HTP-MD dataset used for training different generative models

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2024View details →
zenodo40/100

Hand Washing Video Dataset Annotated According to the World Health Organization's Handwashing Guidelines - METC Subset

<p><strong>Overview:</strong> This is a lab-based dataset with videos recording volunteers (medical students) washing their hands as part of a hand-washing monitoring and feedback experiment. The dataset is collected in the Medical Education Technology Center (METC) of Riga Stradins University, Riga, Latvia. In total, 72 participants took part in the experiments, each washing their hands three times, in a randomized order, going through three different hand-washing feedback approaches (user interfaces of a mobile app). The data was annotated in real time by a human operator, in order to give the experiment participants real-time feedback on their performance. There are 212 hand washing episodes in total, each of which is annotated by a single person. The annotations classify the washing movements according to the World Health Organization&#39;s (WHO) guidelines by marking each frame in each video with a certain movement code.</p> <p>This dataset is part on three dataset series all following the same format:</p> <ul> <li><a href="https://zenodo.org/record/4537209">https://zenodo.org/record/4537209 </a>- data collected in Pauls Stradins Clinical University Hospital</li> <li><a href="https://zenodo.org/record/5808764">https://zenodo.org/record/5808764</a> - data collected in Jurmala Hospital</li> <li><a href="https://zenodo.org/record/5808789">https://zenodo.org/record/5808789</a> - data collected in the&nbsp;Medical Education Technology Center (METC) of Riga Stradins University</li> </ul> <p><strong>Note #1:</strong> we recommend that when using this dataset for machine learning, allowances are made for the reaction speed of the human operator labeling the data. For example, the annotations can be expected to be incorrect a short while after the person in the video switches their washing movements.</p> <p><strong>Application: </strong>The intention of this dataset is to serve as a basis for training machine learning classifiers for automated hand washing movement recognition and quality control.</p> <p><strong>Statistics:</strong></p> <ul> <li>Frame rate: ~16 FPS (slightly variable, as the video are reconstructed from a sequence of jpg images taken with max framerate supported by the capturing devices).</li> <li>Resolution: 640x480</li> <li>Number of videos: 212</li> <li>Number of annotation files: 212</li> </ul> <p>Movement codes (in JSON files):</p> <ul> <li>1: Hand washing movement &mdash; Palm to palm</li> <li>2: Hand washing movement &mdash; Palm over dorsum, fingers interlaced</li> <li>3: Hand washing movement&nbsp;&mdash; Palm to palm, fingers interlaced</li> <li>4: Hand washing movement &mdash; Backs of fingers to opposing palm, fingers interlocked</li> <li>5: Hand washing movement&nbsp;&mdash; Rotational rubbing of the thumb</li> <li>6: Hand washing movement &mdash; Fingertips to palm</li> <li>0: Other hand washing movement</li> </ul> <p><strong>Note #2: </strong>The original dataset of JPG images is available upon request. There are 13 annotation classes in the original dataset: for each of the six washing movements defined by the WHO, &quot;correct&quot; and &quot;incorrect&quot; execution is market with two different labels. In this published dataset, all incorrect executions are marked with code 0, as &quot;other&quot; washing movement.</p> <p><strong>Acknowledgments: </strong>The dataset collection was funded by the Latvian Council of Science project: &quot;Automated hand washing quality control and quality evaluation system with real-time feedback&quot;, No: lzp - Nr. 2020/2-0309.</p> <p><strong>References: </strong>For more detailed information, see this article, describing a similar dataset collected in a different project:</p> <ul> <li> <p>M. Lulla, A. Rutkovskis, A. Slavinska, A. Vilde, A. Gromova, M. Ivanovs, A. Skadins, R. Kadikis, A. Elsts. <em>Hand-Washing Video Dataset Annotated According to the World Health Organization&rsquo;s Hand-Washing Guidelines</em>. Data. 2021; 6(4):38. <a href="https://doi.org/10.3390/data6040038">https://doi.org/10.3390/data6040038</a></p> </li> </ul> <p><strong>Contact information: </strong>atis.elsts@edi.lv</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Hand Washing Video Dataset Annotated According to the World Health Organization's Handwashing Guidelines - Jurmala Hospital Subset

<p><strong>Overview:</strong> This is a large-scale real-world dataset with videos recording medical staff washing their hands as part of their normal job duties in the Jurmala Hospital located in Jurmala, Latvia. There are 2427 hand washing episodes in total, almost all of which are annotated by two persons. The annotations classify the washing movements according to the World Health Organization&#39;s (WHO) guidelines by marking each frame in each video with a certain movement code.</p> <p>This dataset is part on three dataset series all following the same format:</p> <ul> <li><a href="https://zenodo.org/record/4537209">https://zenodo.org/record/4537209</a> - data collected in Pauls Stradins Clinical University Hospital</li> <li><a href="https://zenodo.org/record/5808764">https://zenodo.org/record/5808764</a> - data collected in Jurmala Hospital</li> <li><a href="https://zenodo.org/record/5808789">https://zenodo.org/record/5808789</a> - data collected in the&nbsp;Medical Education Technology Center (METC) of Riga Stradins University</li> </ul> <p><strong>Applications: </strong>The intention of this dataset is twofold: to serve as a basis for training machine learning classifiers for automated hand washing movement recognition and quality control, and to allow to investigate the real-world quality of washing performed by working medical staff.</p> <p><strong>Statistics:</strong></p> <ul> <li>Frame rate: 30 FPS</li> <li>Resolution: 320x240 and 640x480</li> <li>Number of videos: 2427</li> <li>Number of annotation files: 4818</li> </ul> <p>Movement codes (both in CSV and JSON files):</p> <ul> <li>1: Hand washing movement &mdash; Palm to palm</li> <li>2: Hand washing movement &mdash; Palm over dorsum, fingers interlaced</li> <li>3: Hand washing movement&nbsp;&mdash; Palm to palm, fingers interlaced</li> <li>4: Hand washing movement &mdash; Backs of fingers to opposing palm, fingers interlocked</li> <li>5: Hand washing movement&nbsp;&mdash; Rotational rubbing of the thumb</li> <li>6: Hand washing movement &mdash; Fingertips to palm</li> <li>7: Turning off the faucet with a paper towel</li> <li>0: Other hand washing movement</li> </ul> <p><strong>Acknowledgments: </strong>The dataset collection was funded by the Latvian Council of Science project: &quot;Automated hand washing quality control and quality evaluation system with real-time feedback&quot;, No: lzp - Nr. 2020/2-0309.</p> <p><strong>References: </strong>For more detailed information, see this article, describing a similar dataset collected in a different project:</p> <ul> <li> <p>M. Lulla, A. Rutkovskis, A. Slavinska, A. Vilde, A. Gromova, M. Ivanovs, A. Skadins, R. Kadikis, A. Elsts. <em>Hand-Washing Video Dataset Annotated According to the World Health Organization&rsquo;s Hand-Washing Guidelines</em>. Data. 2021; 6(4):38. <a href="https://doi.org/10.3390/data6040038">https://doi.org/10.3390/data6040038</a></p> </li> </ul> <p><strong>Contact information: </strong>atis.elsts@edi.lv</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Training and Test Subsets for Performance Comparison of kNN and GD

<p>The training and test subsets of the&nbsp;fish (https://www.kaggle.com/aungpyaeap/fish-market) and employee (https://www.openml.org/d/42125) dataset. As well as the full employee dataset as CSV.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

A subset dataset of COVID-19 Blood Atlas for CellDrift input

<p>A subset dataset of COVID-19 Blood Atlas for CellDrift input. The original data can be found in this paper:&nbsp;<a href="https://doi.org/10.1016/j.cell.2022.01.012">https://doi.org/10.1016/j.cell.2022.01.012</a>. We did subsetting on the data and extracted 116,124 cells covering 8 disease conditions, 6 PBMC cell types and a series of time points (days since onset) ranging from day 0 to day 25.&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Wikidata subset with revision history information [RDF]

<p>This dataset is composed of 300 instances from the 100 most important classes in Wikidata, for a total of around 30000 entities and 390000 triples. The dataset is geared towards knowledge graph refinement models that leverage edit history information from the graph.&nbsp;There are two versions of the dataset:</p> <ul> <li>The <strong>static</strong> version (files postfixed with &#39;_static&#39;) contains the simple statements of each entity fetched from Wikidata.</li> <li>The <strong>dynamic</strong> version (files postfixed with &#39;_dynamic&#39;) contains information about the operations and revisions made to these entities, and the triples that were added or&nbsp;removed.</li> </ul> <p>Each version is split into three subsets: train, validation (val), and test. Each split contains every entity from the dataset. The train split contains the first 70% of revisions made to each entity, the validation split contains the 70% to 85% revisions, and the test set contains the last 15% revisions.</p> <p>This is a sample from the static datasets:</p> <pre><code>wd:Q217432 a uo:entity ; wdt:P1082 1.005904e+06 ; wdt:P1296 "0052280" ; wdt:P1791 wd:Q18704103 ; wdt:P18 "Pitakwa.jpg" ; wdt:P244 "n80066826" ; wdt:P571 "+1912-00-00T00:00:00Z" ; wdt:P6766 "421180027" .</code></pre> <p>Each entity has the type <em>uo:entity</em>, and contains the statements added during that time period following Wikidata&#39;s data model.</p> <p>In the following code snippet we show an example from the dynamic dataset:</p> <pre><code>uo:rev703872813 a uo:revision ; uo:timestamp "2018-06-28T22:31:32Z" . uo:op703872813_0 a uo:operation ; uo:fromRevision uo:rev703872813 ; uo:newObject wd:Q82955 ; uo:opType uo:add ; uo:revProp wdt:P106 ; uo:revSubject wd:Q6097419 . uo:op703878666_0 a uo:operation ; uo:fromRevision uo:rev703878666 ; uo:opType uo:remove ; uo:prevObject wd:Q1108445 ; uo:revProp wdt:P460 ; uo:revSubject wd:Q1147883 .</code></pre> <p>This dataset is composed of revisions, which have a timestamp. Each revision is composed of 1 to n operations, in which there is a change to a statement from the entity. There are two types of operations: <em>uo:add</em> and <em>uo:remove</em>. In both cases, the property and the subject being modified are shown with the <em>uo:revProp</em> and <em>uo:revSubject</em> properties. In the case of additions, <em>uo:newObject</em> and <em>uo:prevObject</em> properties are added to show the previous and new objects after the addition. In the case of removals, there is a <em>uo:prevObject </em>property to record the object that was removed.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Figure 3: Percentages of women in the non-atypical academic workforce in a subset of higher education institutions in the United Kingdom by institutional groupings, 2021

<p>This Zenodo entry includes the full data files (.csv and .xlsx)&nbsp;for Figure 3: <em>Percentages of women in the non-atypical academic workforce in a subset of higher education institutions in the United Kingdom by institutional groupings, 2021</em>&#39;,&nbsp;included in the manuscript &quot;The Curtin Open Knowledge Initiative: Sharing data on scholarly research performance&quot;, authored by members of the Curtin Open Knowledge Initiative (COKI). The&nbsp;analysis is of&nbsp;publicly available data sourced from the United Kingdom Higher Education Statistics Agency (HESA).</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Subset de la colección de notas de La Nación

<p>Se trata de un subset de notas period&iacute;sticas del diario argentino La Naci&oacute;n que cubre desde el 2016 hasta 2019. Contiene las siguientes variables: fecha,&nbsp;titulo, nota.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

SDSS Galaxy Subset

<p>The&nbsp;<a href="https://www.sdss.org">Sloan Digital Sky Survey</a>&nbsp; (SDSS) is a comprehensive survey of the northern sky. This dataset contains a subset of this survey, of 100077&nbsp;objects classified as galaxies, it includes a CSV file with a collection of information and a set of files for each object, namely JPG image files, FITS and spectra data. This dataset is used to train and explore the&nbsp;<a href="https://github.com/nunorc/astromlp-models">astromlp-models</a>&nbsp;collection of deep learning models for galaxies characterisation.</p> <p>The dataset includes a CSV data file where each row is an object from the SDSS database, and with the following columns (note that some data may not be available for all objects):</p> <ul> <li><strong>objid</strong>: unique SDSS object identifier&nbsp;</li> <li><strong>mjd</strong>: MJD of observation</li> <li><strong>plate</strong>: plate identifier</li> <li><strong>tile</strong>: tile identifier</li> <li><strong>fiberid</strong>: fiber identifier</li> <li><strong>run</strong>: run number</li> <li><strong>rerun</strong>: rerun number</li> <li><strong>camcol</strong>: camera column</li> <li><strong>field</strong>: field number</li> <li><strong>ra</strong>: right ascension</li> <li><strong>dec</strong>: declination</li> <li><strong>class</strong>: spectroscopic class (only objetcs with GALAXY are included)</li> <li><strong>subclass</strong>: spectroscopic subclass</li> <li><strong>modelMag_u</strong>: better of DeV/Exp magnitude fit for band u</li> <li><strong>modelMag_g</strong>: better of DeV/Exp magnitude fit for band g</li> <li><strong>modelMag_r</strong>: better of DeV/Exp magnitude fit for band r</li> <li><strong>modelMag_i</strong>: better of DeV/Exp magnitude fit for band i</li> <li><strong>modelMag_z</strong>: better of DeV/Exp magnitude fit for band z</li> <li><strong>redshift</strong>: final redshift from SDSS data z</li> <li><strong>stellarmass</strong>: stellar mass extracted from the&nbsp;<a href="https://www.sdss.org/dr16/spectro/eboss-firefly-value-added-catalog">eBOSS Firefly catalog</a></li> <li><strong>w1mag</strong>: WISE W1 &quot;standard&quot; aperture magnitude</li> <li><strong>w2mag</strong>: WISE W2 &quot;standard&quot; aperture magnitude</li> <li><strong>w3mag</strong>: WISE W3 &quot;standard&quot; aperture magnitude</li> <li><strong>w4mag</strong>: WISE W4 &quot;standard&quot; aperture magnitude</li> <li><strong>gz2c_f</strong>: Galaxy Zoo 2 classification from&nbsp;<a href="https://academic.oup.com/mnras/article/435/4/2835/1022913">Willett et al 2013</a></li> <li><strong>gz2c_s</strong>: simplified version of Galaxy Zoo 2 classification (<a href="https://github.com/nunorc/astromlp-models#galaxy-zoo-2-simplified-classes-gz2c">labels set</a>)</li> </ul> <p>Besides the CSV file a set of directories are included in the dataset, in each directory you&#39;ll find a list of files named after the <strong>objid&nbsp;</strong>column from the CSV file, with the corresponding data, the following directories tree is&nbsp;available:</p> <pre><code class="language-bash">sdss-gs/ ├── data.csv ├── fits ├── img ├── spectra └── ssel</code></pre> <p>Where, each directory contains:</p> <ul> <li><strong>img</strong>: RGB images from the object in JPEG format, 150x150 pixels, generated using the&nbsp;<a href="https://skyserver.sdss.org/dr16/en/help/docs/api.aspx">SkyServer DR16 API</a></li> <li><strong>fits</strong>: FITS data subsets around the object across the u, g, r, i, z bands; cut is done using the&nbsp;<a href="https://github.com/jhoar/ImageCutter">ImageCutter</a>&nbsp;library</li> <li><strong>spectra</strong>: full best fit spectra data from SDSS between 4000 and 9000 wavelengths</li> <li><strong>ssel</strong>: best fit spectra data from SDSS for specific selected intervals of wavelengths discussed by&nbsp;<a href="https://arxiv.org/abs/1003.3186">S&aacute;nchez Almeida 2010</a></li> </ul> <p><strong>Changelog</strong></p> <ul> <li>v0.0.4&nbsp;- Increase number of&nbsp;objects to ~100k.</li> <li>v0.0.3&nbsp;- Increase number of&nbsp;objects to ~80k.</li> <li>v0.0.2&nbsp;- Increase number of&nbsp;objects to ~60k.</li> <li>v0.0.1 - Initial import.</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

LC-MS² meta data for each MassBank (MB) subset

<p>The CSV-file (tab used as separator) provides the Liquid-chromatography (LC) and Tandem-mass spectrometry (MS&sup2;) configurations for each MassBank (MB) subset used in the publication: &quot;Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data&quot; by Bach et al. (2022).</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Savi et al., 2020 -- tributary-main-channel interaction experiments -- Experiment No Change 2 subset as netCDF files

<p><strong>Overview</strong></p> <p>Zip file contains two netCDF files with a subset of&nbsp;data from the &quot;No Change 2&quot; (NC2) experiment conducted by Savi et al., 2020 and published in Earth Surface Dynamics (<a href="https://doi.org/10.5194/esurf-8-303-2020">https://doi.org/10.5194/esurf-8-303-2020</a>) with the original data available via the Sediment Experimentalists Network Project Space SEAD Internal Repository (<a href="https://doi.org/10.26009/s0ZOQ0S6">https://doi.org/10.26009/s0ZOQ0S6</a>). Topographic scan data were&nbsp;re-formatted into the netCDF file &quot;T_NC2_scans.nc&quot;, and overhead imagery was extracted from the video of the experiment&nbsp;approximately once every minute of experimental time and RGB band data is provided in the formatted netCDF file &quot;T_NC2_images.nc&quot;. These data were formatted into netCDF files for easy loading into the &quot;deltametrics&quot; analysis toolbox.</p> <p>&nbsp;</p> <p><strong>Additional Details</strong></p> <p>Re-packaging the scan data from the .tif files was straightforward. From the metadata&nbsp;spreadsheet, we know the times at which the scans were taken (and can eliminate the redundant scan). From the paper itself we know the resolution of the topographic scans is 1 mm in the horizontal and vertical. We also know the input discharges, both water and sediment, through both the main channel and tributary, from the paper. We provide these values as metadata in the netCDF files. The scans form the &#39;eta&#39; field representing the topography in the file. The packaged up netCDF file is called &#39;T_NC2_scans.nc&#39;.</p> <p>Overhead imagery&nbsp;from the T_NC2_Complete21fps.wmv video file was extracted using the following command:</p> <blockquote> <p>ffmpeg -i T_NC2_Complete21fps.wmv -r 21 T_NC2_frames/%04d.png</p> </blockquote> <p>This command utilizes the ffmpeg tool to extract the frames at a rate of 21 frames per second (-r 21) as the file name implies that is the rate at which the overhead photos were combined into a video. The NC designation indicates that this experiment was performed with no change in the input conditions in either the main or tributary channels.</p> <p>The experiment ran for a total of 480 minutes. A total of 1466 images were obtained from the ffmpeg extraction. This translates to an image approximately every 20 seconds of real time (480 minutes / 1466 frames * 60 seconds/minute = 19.6453 seconds / frame). We sample every 3rd frame, which gives us images roughly once a minute (489 frames in all), to create the subset of data re-packaged as a netCDF file for deltametrics. Dimensions for the pixels were approximated based on our knowledge of the topographic scan resolution. Assuming the extents of the scans and overhead images are the same (although they are not), we calculate the number of millimeters per pixel in the x and y directions for the overhead images. We assume the pixels are more likely to be square than rectangular, so we average these values and assign this as the distance per pixel in both the x and y dimensions for these data.</p> <p>Script used to re-package this dataset is available as a <a href="https://gist.github.com/elbeejay/f7603e712cf0ea30a8215b902715d222">GitHub Gist</a>.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>Savi, Sara, et al. &quot;Interactions between main channels and tributary alluvial fans: channel adjustments and sediment-signal propagation.&quot; Earth Surface Dynamics 8.2 (2020): 303-322.</p> <p>Physical experiments on interactions between main-channels and tributary alluvial fans<br> S. Savi, Tofelde, A. Wickert, A. Bufe, T. Schildgen, and M. Strecker<br> https://doi.org/10.26009/s0ZOQ0S6</p>

opencc-bySep 2022View details →
zenodo40/100

Subset of 300 out of 3000 Prepared Sentinel 2 Scenes for Transfer Learning and Super-sampling.

<p>This dataset contains a random subset of 300 out of 3000 Sentinel 2 scenes prepared for Transfer Learning and Super-sampling. It comes in the form of zipped NumPy arrays in the npz format. The files have are named in the following fashion:</p> <p>latutide+latitude_decimals_longitude+longitude_decimals_month_of_the_year_for_mosaic.&nbsp;</p> <p>Each file contains:</p> <p>bands.npy: The Sentinel 2 bands in 10m resolution uint10: B02, B03, B04, B08, B05, B06, B07, B8A, B11, B12. The 20m bands have been resampled using bilinear resampling.</p> <p>nir.npy: B08 Resampled to 20m using average resampling and then resampled to 10m using bilinear. Useful for training super-sampling models.</p> <p>scl.npy: The Sentinel 2 Scene Classification file. Contains information on cloud cover and land cover.</p> <p>sincos.npy: Contains the latitude, longitude, and time of capture for each pixel encoded to sine and cosine waves in the [0,1] interval. The is useful when training a model to predict where on the globe an image was captured.</p> <p>The images can be processed to patches using the buteo toolbox:&nbsp;</p> <p>`pip install buteo --upgrade</p> <p>`import buteo as beo`</p> <p>`beo.get_patches(beo.raster_to_array(&quot;path_to_bands&quot;))`</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record