Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,863

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,863 results for “Challenge”

Learn how ShareScore rates datasets ↗
zenodo44/100

ICDAR'15 SMARTPHONE DOCUMENT CAPTURE AND OCR COMPETITION (SmartDoc) - Challenge 1 (original version)

<p><strong>CHALLENGE 1: SMARTPHONE DOCUMENT CAPTURE COMPETITION</strong></p> <p><strong>Smartphones are replacing personal scanners.</strong>&nbsp;They are portable, connected, powerful and affordable. They are on their way to become the new entry point in business processing applications like document archival, ID scanning, check digitization, just to name a few. In order keep our workflows streamlined,&nbsp;<strong>we need to make those new capture device as reliable as batch scanners</strong>.</p> <p>We believe an efficient capture process should be able to:</p> <ol> <li><em>detect and segment</em>&nbsp;the relevant document object during the preview phase;</li> <li><em>assess the quality</em>&nbsp;of the capture conditions and help the user improve them;</li> <li>optionally,&nbsp;<em>trigger the capture</em>&nbsp;at the perfect moment;</li> <li>and&nbsp;<em>produce a high-quality, controlled output</em>&nbsp;based on the high resolution captured image.</li> </ol> <p>This competition is focused on the first step of this process:<strong>&nbsp;</strong><strong>efficiently detect and segment document regions</strong>, as illustrated by following video showing the ideal output for the preview phase of some acquisition session: <a href="https://youtu.be/WNsI0R_rpO0">Click here to watch the video.</a> This video shows the ideal document object detection &lrm;(ie the ground truth, as a red frame)&lrm;.</p> <p>For this challenge, the <strong>input</strong> consists in a set of <strong>videoclips containing a document</strong> from a predefined set, and the <strong>output</strong> should be an <strong>xml file containing the quadrilateral coordinates</strong> in which we can find the document per each frame of the video. Click <a href="https://sites.google.com/site/icdar15smartdoc/challenge-1/challenge1dataset">here</a> for detailed information about the dataset.&nbsp;</p> <p>&nbsp;</p> <p><strong>Licence</strong> for the dataset of challenge 1 (page outline detection in preview frames) :</p> <p>This work is licensed under a <strong>Creative Commons Attribution 4.0 International License</strong> &lt;<a href="https://www.google.com/url?q=http://creativecommons.org/licenses/by/4.0/&amp;sa=D&amp;ust=1524734857667000&amp;usg=AFQjCNEt4YXnUv2nXCFwkeOuBDqxDpvknQ">http://creativecommons.org/licenses/by/4.0/</a>&gt;. Author attribution should be given by citing the following conference paper: Jean-Christophe Burie, Joseph Chazalon, Micka&euml;l Coustaty, S&eacute;bastien Eskenazi, Muhammad Muzzamil Luqman, Maroua Mehri, Nibal Nayef, Jean-Marc OGIER, Sophea Prum and Mar&ccedil;al Rusinol: &ldquo;ICDAR2015 Competition on Smartphone Document Capture and OCR (SmartDoc)&rdquo;, In 13th International Conference on Document Analysis and Recognition (ICDAR), 2015.</p> <p><strong>If you use this dataset, please send us a short email at &lt;icdar.smartdoc (at) gmail.com&gt; to tell us why it was useful to you, and whether you have results or publications we can reference on our website. Thank you!</strong></p>

opencc-by-4.0Aug 2015View details →
zenodo44/100

Dataset for "Contrasting Effects of Organic and Mineral Nitrogen Challenge the N-Mining Hypothesis for Soil Organic Matter Priming"

<p>Dataset for the article:</p> <p>Mason-Jones, K., Schm&uuml;cker, N., Kuzyakov, Y. (2018) Contrasting Effects of Organic and Mineral Nitrogen Challenge the N-Mining Hypothesis for Soil Organic Matter Priming. Soil Biology and Biochemistry 124, 38-46, https://doi.org/10.1016/j.soilbio.2018.05.024</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

MICCAI 2016 MS lesion segmentation challenge: supplementary results

<p>This package contains supplementary material for our article prepared for publication and under revision. It contains omitted results due to space limits of the article as well as detailed, patient per patient and team per team results for all metrics. Additional figures redundant with those of the article are also provided.&nbsp;</p> <p>The readme file Readme_SupplementalMaterial.txt provides details about each individual file content.</p>

opencc-by-4.0Jul 2018View details →
zenodo44/100

IML workshop challenge on jet mass regression

<p>This dataset is associated with the LPCC IML (Lhc Physics Center at Cern Inter-experimental Machine Learning) working group.&nbsp; It was produced for the second IML annual workshop (April 2018).</p> <p>This dataset is part of a machine learning &quot;challenge&quot; on jet mass regression at future circular collider (FCC) conditions.&nbsp; Further details can be found on the challenge page, here:</p> <p><a href="https://gitlab.cern.ch/IML-WG/IML_challenge_2018/wikis/home">https://gitlab.cern.ch/IML-WG/IML_challenge_2018/wikis/home</a></p>

opencc-zeroMar 2018View details →
zenodo44/100

CRT-EPiggy19 challenge

<p><em><strong>OBJECTIVE OF THE CHALLENGE</strong></em></p> <p>The spirit of CRT-EPiggy19 is to <strong>collectively review the current state-of-the-art for computational cardiology models and their ability to predict pacing-based therapy outcomes</strong>, as well as the <strong>identification of the most critical phases and more promising solutions in the personalization modelling pipeline</strong>.</p> <p>More specifically, participants will be asked to <strong>predict the electrical response of CRT</strong> and to <strong>propose the optimal device configuration</strong> in a swine model of left bundle branch block, given fully controlled data. All challenge participants will be invited to contribute to the preparation of a journal article summarizing the main findings from the CRT-EPiggy19 challenge, similarly to the CESC&rsquo;10 challenge (<a href="https://www.sciencedirect.com/science/article/pii/S0079610711000708">Camara et al., Prog Biophys Mol Bio 2011</a>).</p> <p><em><strong>TRAINING/TEST DATASETS</strong></em></p> <p>Some years ago, researchers at Hospital Cl&iacute;nic de Barcelona and Universitat Pompeu Fabra developed a swine model of left bundle branch block (LBBB) for experimental studies of CRT (<a href="https://link.springer.com/article/10.1007%2Fs12265-013-9464-1">Rigol et al., J Cardiovasc Transl Res 2013</a>). Radiofrequency applications were performed to induce LBBB, and half of the animals presented a myocardial infarction located at the septal wall. Imaging data and electro-anatomical maps (EAM) were acquired at baseline, with the induced LBBB and after implantation of a CRT devicE. This rich data is well suited for evaluating some features of the different cardiac computational models available nowadays, and will be the basis of the CRT-EPiggy19 challenge. The training data will include two complete infarcted and two non-infarcted datasets (total of 4 cases), while the test data is composed of four cases for each of the two categories (infarcted vs non-infarcted; total of 8 cases). The electrical activation patterns of the training datasets have already been described with detail in <a href="https://ieeexplore.ieee.org/document/7786893">Soto-Iglesias et al. (IEEE J Transl Eng Health Med 2016)</a>. Check the Datasets section for a preview of the training data and the procedure to download it. Unlike LBBB and CRT activation maps, baseline maps will not initially be released, since they do not necessarily contribute to the prediction of CRT from LBBB.</p> <p>Some of the main sources of variability in the personalization of cardiac models involve the extraction of anatomical data from medical images and the creation of the geometrical domain where models are run. In order to reduce this variability in the CRT-EPiggy19 challenge, biventricular finite element meshes will be provided to each participant, which were built from the segmentation of cine-MRI data. These meshes will include cardiomyocyte orientation (obtained with rule-based models; see <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/cnm.3185">Doste et al. Int J Numer Meth Bio 2019</a> for details), several regional labels (AHA regions, endo- and epi-cardial walls, different ventricles) and the local activation times projected from EAM data. Additionally, the affected AHA segments and its transmurality will be given for infarcted cases. Furthermore, for visualization and analysis purposes, 2D bi-ventricular representations will be given.</p> <p><em><strong>EVALUATION METRICS</strong></em></p> <p>Global and regional differences between simulated and measured CRT activation maps will be used to evaluate the prediction accuracy of each proposed model. As global metric, we will use the <strong>difference in Total Activation Time</strong> (TAT, whole heart fully activated). The TAT will also individually be assessed for the LV, the RV, as well as for each AHA segment. TAT differences will be separately analysed between simulations and measurements for infarcted vs. non-infarcted cases. <strong>Histograms of isochrones of electrical activation</strong> will be derived from simulations to estimate <strong>inter- and intra-ventricular electrical dyssynchrony</strong> (<a href="https://ieeexplore.ieee.org/document/7786893">Soto-Iglesias et al., IEEE J Transl Eng Health Med 2016</a>). Each participant will be asked to report the <strong>used hardware infrastructure, computational times and details about the implementation</strong> and a <strong>self-reported analysis for model integration onto a clinical workflow</strong>.</p> <p><em><strong>CONTACT</strong></em></p> <p>You are welcome to contact Oscar Camara should you have any questions at: oscar.camara@upf.edu. More details on the CRT-EPiggy19 challenge can be found in the following website:&nbsp;<a href="http://crt-epiggy19.surge.sh/">crt-epiggy19.surge.sh</a>.</p> <p><em><strong>LICENSE</strong></em></p> <p>All data at the CRT-EPiggy19 challenge are released under Creative Commons (CC) licenses.&nbsp;</p> <p><em><strong>CITATION</strong></em></p> <p>If you use the CRT-EPiggy19 challenge dataset or part of it, please cite the following paper:&nbsp;<a href="https://link.springer.com/article/10.1007%2Fs12265-013-9464-1">Rigol et al., J Cardiovasc Transl Res 2013</a>.</p> <p>Rigol M, Solanes N, Fernandez-Armenta J, Silva E, Doltra A, Duchateau N, Barcelo A, Gabrielli L, Bijnens B, Berruezo A, Brugada J, Sitges M. Development of a swine model of left bundle branch block for experimental studies of cardiac resynchronization therapy. J&nbsp;Cardiovasc Transl Res. 2013 Aug;6(4):616-22. doi: 10.1007/s12265-013-9464-1.</p> <p>You may also consider citing the paper, which describes the electrophysiological patterns of the training data:&nbsp;<a href="https://ieeexplore.ieee.org/document/7786893">Soto-Iglesias et al., IEEE J Transl Eng Health Med 2016</a></p> <p>Soto Iglesias D, Duchateau N, Kostantyn Butakov CB, Andreu D, Fernandez-Armenta J, Bijnens B, Berruezo A, Sitges M, Camara O. Quantitative Analysis of Electro-Anatomical Maps: Application to an Experimental Model of Left Bundle Branch Block/Cardiac Resynchronization Therapy. IEEE J Transl Eng Health Med. 2016 Dec 16;5:1900215. doi: 10.1109/JTEHM.2016.2634006.</p> <p><em><strong>ACKNOWLEDGMENTS</strong></em></p> <p>This work was supported in part by the Spanish Ministry of Science and Innovation (TIN2011-28067, REDINSCOR RD06/003/008), the Spanish Industrial and Technological Development Center (cvREMOD-CEN-20091044), the Seventh Framework Programme (FP7/2007-2013) for research, technological and demonstration under grant agreement VP2HF (no. 611823), and the Spanish Ministry of Economy and Competitiveness under the Maria de Maeztu Units of Excellence Programme (MDM-2015-0502).</p>

opencc-by-4.0Apr 2019View details →
zenodo44/100

The Consonant Challenge Corpus

<p>The Consonant Challenge Corpus provides a dataset to support&nbsp;human-machine comparisons of consonant recognition in quiet and noise.&nbsp;Twelve female and 12 male native English talkers contributed to the corpus. All speakers produced each of the 24 English consonants / b, d, g, p, t, k, s, ʃ, f, v, &eth;, &theta;, ʧ, z, ʒ, h, ʤ, m, n, ŋ, w, r, j, l / in nine vowel contexts consisting of all possible combinations of the three vowels / iː / (as in &ldquo;beat&rdquo;), / uː / (as in &ldquo;boot&rdquo;), and / &aelig; / (as in &ldquo;bat&rdquo;). Each VCV was produced using both front and end stress (e.g. / &lsquo;&aelig; b &aelig; / vs / &aelig; b &lsquo;&aelig; /) giving a total of 24 (speakers) * 24 (consonants) * 2 (stress types) * 9 (vowel contexts) = 10368 tokens. Tokens are distributed into training, development and test sets for the purposes of automatic speech recognition experiments.</p> <p>The Consonant Challenge is described in this article:&nbsp;Cooke, M., Scharenborg, O. (2008), &ldquo;The Interspeech 2008 Consonant Challenge&rdquo;, Proceedings of Interspeech, Brisbane, Australia, September 2008.</p> <p>The&nbsp;distribution consists of the following elements:&nbsp;</p> <p>Technical description:</p> <ul> <li><em>readme</em></li> </ul> <p>Speech/noise waveforms</p> <ul> <li><em>train.zip</em> contains noise-free training data</li> <li><em>test.zip</em> contains the 7 test sets as well as practice items for perceptual tests, and MATLAB format files containing offsets identifying the time location of the speech token within the mixture</li> <li><em>test_binaural.zip</em> contains 2-channel wavs with the speech and noise on separate channels (left=noise, right=speech), for test sets 2-7 (test set 1 is noise-free)</li> <li><em>dev.zip</em> development set</li> <li><em>dev_binaural.zip</em>&nbsp;is the 2-channel version of the development set</li> </ul> <p>Phoneme segmentation data</p> <ul> <li><em>handsegm.91.mlf.txt:</em> 91 hand-segmented VCVs in HTK format.&nbsp;This set consists of at least three items per consonant in a context in which the first and the second vowel were identical, added to that were 19 randomly selected VCVs.</li> <li><em>segmentation_training.mlf.txt:</em> automatically generated phoneme segmentation of the clean training material&nbsp;in HTK format</li> <li><em>segmentation_testsets.zip</em>: zip file containing automatically generated phoneme segmentations of each test set&nbsp;in HTK format&nbsp;</li> </ul> <p>Automatic speech recognition</p> <ul> <li><em>asr.zip</em> contais&nbsp;scripts and models</li> </ul>

opencc-by-4.0Nov 2019View details →
zenodo44/100

Visual abstract for SAMPL6 logP Challenge

<p>This figure was created as a visual abstract for the SAMPL6 Part II logP Challenge, which was a blind computational prediction challenge for predicting octanol-water partition coefficients of kinase inhibitor fragment-like small molecules. This figure&nbsp;can possibly be used as a cover art for the special journal issue organized for this&nbsp;challenge.&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

Cadenza Challenge (CAD1): databases for the First Cadenza Challenge - Task1

<h2>Cadenza</h2> <p>This is the training, validation and evaluation data for the&nbsp;<a href="https://cadenzachallenge.org/docs/cadenza1/cc1_intro" target="_blank" rel="noopener">First Cadenza Challenge - Task 1</a>.</p> <p>The Cadenza Challenges are improving music production and processing for people with a hearing loss. According to The World Health Organization, 430 million people worldwide have a disabling hearing loss. Studies show that not being able to understand lyrics is an important problem to tackle for those with hearing loss. Consequently, this task is about improving the intelligibility of lyrics when listening to pop/rock over headphones. But this needs to be done without losing too much audio quality - you can't improve intelligibility just by turning off the rest of the band! We will be using one metric for intelligibility and another metric for audio quality, and giving you different targets to explore the balance between these metrics.</p> <p>Please see the&nbsp;<a href="https://cadenzachallenge.org/">Cadenza website</a> for a full description of the data</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Cadenza Challenge (ICASSP24): databases for ICASSP 2024 Cadenza Grand Challenge

<h2>Cadenza</h2> <p>This is the training, validation and evaluation data for the <a href="https://cadenzachallenge.org/docs/icassp_2024/intro" target="_blank" rel="noopener">ICASSP 2024 Cadenza Grand Challenge</a>.</p> <p>The Cadenza Challenges are improving music production and processing for people with a hearing loss. According to The World Health Organization, 430 million people worldwide have a disabling hearing loss. Studies show that not being able to understand lyrics is an important problem to tackle for those with hearing loss. Consequently, this task is about improving the intelligibility of lyrics when listening to pop/rock over headphones. But this needs to be done without losing too much audio quality - you can't improve intelligibility just by turning off the rest of the band! We will be using one metric for intelligibility and another metric for audio quality, and giving you different targets to explore the balance between these metrics.</p> <p>Please see the&nbsp;<a href="https://cadenzachallenge.org/">Cadenza website</a> for a full description of the data</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Supplementary data: Breeding wheat for organic farming: can the high grain protein gene Gpc-B1 help to tackle challenges in view of end-use quality?

<p>Agronomic and quality data of organic wheat (<em>Triticum aestivum</em>), mean comparisons and supplementary figures related to the publication "Breeding wheat for organic farming: can the high grain protein gene Gpc-B1 help to tackle challenges in view of end-use quality?" by Grausgruber et al. (2024) published in the Journal of Cereal Science.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Cadenza Challenge (CAD1): databases for the First Cadenza Challenge - Task2

<p>This is the training, validation and evaluation data for the&nbsp;<a href="https://cadenzachallenge.org/docs/cadenza1/cc1_intro" target="_blank" rel="noopener">First Cadenza Challenge - Task 2</a> listening music in a car.</p> <p>The Cadenza Challenges are improving music production and processing for people with a hearing loss. According to The World Health Organization, 430 million people worldwide have a disabling hearing loss. Studies show that not being able to understand lyrics is an important problem to tackle for those with hearing loss. Consequently, this task is about improving the intelligibility of lyrics when listening to pop/rock over headphones. But this needs to be done without losing too much audio quality - you can't improve intelligibility just by turning off the rest of the band! We will be using one metric for intelligibility and another metric for audio quality, and giving you different targets to explore the balance between these metrics.</p> <p>Please see the&nbsp;<a href="https://cadenzachallenge.org/">Cadenza website</a> for a full description of the data</p> <h3>File description</h3> <ul> <li><strong>cadenza_cad1_task2_core.v1_1.tar.gz</strong>: Audio and metadata files for training and validation.</li> <li><strong>cadenza_cad1_task2_evaluation.v1_1.tar.gz</strong>: Audio and metadata files for evaluation</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Supplementary material for journal article "Challenge dose titration in a Mycobacterium bovis infection model in goats"

<p>Supplementary Figure and Table to Journal article. Figure shows daily rectal temperature of each animal after inoculation. Table 1 shows number and volume of pulmonary lesions for each animal as detected by computed tomography imaging. Table 2 gives details about scoring used at clinical examination.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Ressources for End-to-End French Text-to-Speech Blizzard challenge

<p>Here are 289 chapters of 5 audiobooks from Librivox (51:12) read by Nadine Eckert-Boulet (NEB):</p> <ol> <li>Madame Bovary (MB) by Gustave Flaubert (FL) - 3 volumes, 35 chapters<br>(original <a href="https://librivox.org/madame-bovary-french-by-gustave-flaubert">wavs</a>; <a href="https://www.gutenberg.org/cache/epub/14155/pg14155.txt">text</a>)</li> <li>Les myst&egrave;res de Paris (LMP) by Eugene Sue (ES) - 4 volumes, 83 chapters (original <a href="https://librivox.org/les-mysteres-de-paris-tome-1-by-eugene-sue">wavs1</a>, <a href="https://librivox.org/les-mysteres-de-paris-tome-2-by-eugene-sue/">wavs2</a>,<a href="https://librivox.org/les-mysteres-de-paris-tome-3-by-eugene-sue"> wavs3</a>; <a href="https://www.gutenberg.org/cache/epub/18921/pg18921.txt">text1</a>, <a href="https://www.gutenberg.org/cache/epub/18922/pg18922.txt">text2</a>, <a href="https://www.gutenberg.org/cache/epub/18923/pg18923.txt">text3</a>)</li> <li>Les tribulations d'un chinois en Chine (TCC) by Jules Verne (JV) - 1 volume, 22 chapters (original <a href="https://librivox.org/les-tribulations-dun-chinois-en-chine-by-jules-verne">wavs</a>; <a href="https://www.gutenberg.org/cache/epub/14162/pg14162.txt">text</a>)</li> <li>La fille du pirate (LFDP) by Henri &Eacute;mile Chevalier (EC) - 7 volumes, 121 chapters (original <a href="https://librivox.org/la-fille-du-pirate-by-henri-emile-chevalier">wavs</a>, <a href="https://www.gutenberg.org/cache/epub/18403/pg18403.txt">text)</a></li> <li>La vampire (VAMP) by Paul F&eacute;val (PF) - 1 volume, 28 chapters (original <a href="https://librivox.org/la-vampire-by-feval-paul-henry-corentin">wavs</a>, <a href="https://www.gutenberg.org/cache/epub/10053/pg10053.txt">text</a>)</li> </ol> <p>and</p> <p>2515 utterances (2:03) read by another female French speaker Aur&eacute;lie Derbier (AD):</p> <ol> <li>1608 utterances extracted from various books (DIVERS_BOOK_AD*)</li> <li>907 transcripts of the sessions of the French parliament (DIVERS_PARL_01*)</li> </ol> <p>We recently added three speakers from Librivox/Litteratureaudio:</p> <ol> <li>Ezwa (EZWA): L'&eacute;pouvante by Maurice Level (original&nbsp;<a href="https://librivox.org/lepouvante-by-maurice-level-1010/">wavs</a>; <a href="https://www.gutenberg.org/cache/epub/17794/pg17794.txt">text</a>) - 11 chapters - 4869 utterances&gt; 03:16</li> <li>Pauline Latournerie (PL): Le p&eacute;dagogue n'aime pas les enfants by Henri Roorda (<a href="https://librivox.org/le-pedagogue-naime-pas-les-enfants-by-henri-roorda/">original wavs</a>; <a href="https://ebooks-bnr.com/ebooks/pdf4/roorda_le_pedagogue_n_aime_pas_les_enfants.pdf">text</a>) - 6 chapters - 1320 utterances&gt; 01:17</li> <li>Jean-Luc Fischer (JLF): L&rsquo;Affaire Charles Dexter Ward by Howard Phillips Lovecraft (<a href="https://www.litteratureaudio.com/livre-audio-gratuit-mp3/howard-phillips-lovecraft-laffaire-charles-dexter-ward.html">original wavs</a>; <a href="https://www.litteratureaudio.com/textes/H_P_Lovecraft_L_Affaire_CDW.pdf">text</a>) - 16 chapters - 1823 utterances&gt; 02:37</li> </ol> <p>Each .wav file (sampled at 22050Hz) corresponds to one entire chapter. The format of the filenames is:<br>{author's acronym}_{book's acronym}_{reader's acronym}_{volume's number}_{chapter's number}</p> <p>The NEB_train.csv file gives text and phonetic alignments (essentially for MB and LMP) for utterances in 4 fields separated by '|':<br>{filename}|{start_ms}|{end_ms}|{text or phonetic content}. Most utterances are separated by at least a pause of 400ms. The intervals [start_ms:end_ms] comprise leading and trailing silences of 130ms (since wavs are entire chapters, these silences are "true" ambient silences). Same for AD_train.csv.</p> <p>When phonetic alignment has been performed, 2 additional fields have been added: {aligned phones}|{durations in ms}. Each input character or phone has a corresponding aligned phone and a duration. Note that all aligned utterances start and end with an aligned phone of 130ms. The set of aligned phones comprises:</p> <ul> <li>The set of input phones</li> <li>The silence: '__'</li> <li>The symbol&nbsp;'_'&nbsp;for silent characters, e.g. "chat" is aligned with&nbsp;'s^ _ a _'</li> <li>29 combined&nbsp;aligned phones ('a&amp;i', 'a&amp;j', 'b&amp;q', 'd&amp;q','d&amp;z', 'd&amp;z^', 'f&amp;q', 'g&amp;q', 'g&amp;z', 'j&amp;i', 'j&amp;u', 'j&amp;q', 'i&amp;j', 'k&amp;q', 'k&amp;s', 'k&amp;s&amp;q', 'l&amp;q', 'm&amp;q', 'n&amp;q',&nbsp;'r&amp;w', 'r&amp;q', 's&amp;q', 't&amp;q', 't&amp;s', 't&amp;s^', 'w&amp;a', 'z&amp;q', 'p&amp;q') that align to only one&nbsp;character,&nbsp;e.g. "expatrier" is aligned with&nbsp;'e^ k&amp;s p a t&nbsp;r i&amp;j e _'</li> </ul> <p>Text is in UTF8. '&laquo;&raquo;','&not;', '~','""','()','[]' are respectively used for speaking quotes, turn switches, three dots, quoted expression, aside quotes, notes. Because of rare occurrences, '&ouml;' has been transcribed as 'oe'. Paragraphs (two consecutive carriage returns in the original text) are cued by a special character '&sect;'. It usually ends an utterance but could be used within an utterance if its associated pause is too short.</p> <p>When available, phonetic content is given per word in curly brackets '{}'. We use 39 phonetic symbols:</p> <ul> <li><strong>oral vowels</strong>: a (f<strong><em>a</em></strong>), e (f<em><strong>&eacute;e</strong></em>), e^ (f<em><strong>ait</strong></em>), x (f<em><strong>eu</strong></em>), x^ (c<em><strong>oeu</strong></em>r), i (r<em><strong>iz</strong></em>), y (f<em><strong>ut</strong></em>), u (f<em><strong>ou</strong></em>), o (f<em><strong>aux</strong></em>), o^ (p<strong><em>o</em></strong>rc)</li> <li><strong>schwa</strong>: q (gag<strong><em>e</em></strong>)</li> <li><strong>nasal vowels</strong>: a~ (r<strong><em>an</em></strong>g), e~ (f<em><strong>in</strong></em>), x~ (<strong><em>un</em></strong>), o~ (r<em><strong>on</strong></em>d)</li> <li><strong>semi-vowels</strong>: h (h<em><strong>u</strong></em>it), w (<strong><em>ou</em></strong>ate), j (h<em><strong>i</strong></em>er)</li> <li><strong>consonants</strong>: p (<em><strong>p</strong></em>as), t (<strong><em>t</em></strong>as), k (<em><strong>c</strong></em>as), b (<strong><em>b</em></strong>as), d (<em><strong>d</strong></em>os), g (<em><strong>g</strong></em>ars), f (<em><strong>f</strong></em>aux), s (<strong><em>s</em></strong>ot) , s^ (<strong><em>ch</em></strong>at), v (<strong><em>v</em></strong>u), z (<strong><em>z</em></strong>ut), z^ (<em><strong>j</strong></em>us), r (<strong><em>r</em></strong>iz), l (<em><strong>l</strong></em>a), m (<strong><em>m</em></strong>a), n (<strong><em>n</em></strong>on), n~ (oi<strong><em>gn</em></strong>on), ng (campi<em><strong>ng</strong></em>)</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Supplementary data to Challenges and Opportunities for the Recovery of Critical Raw Materials from Electronic Waste

<p>Content of the excell file:</p> <p>Table S1: Relationship between UNU-keys and&nbsp;MINCOTUR codes</p> <p>Table S2: UNU-Key composition and alpha and beta values for Weibull distributions</p> <p>Table S3: Metal prices used in this study.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Codes and Data for 'Cost-effective Planning of Decarbonized Power-Gas Infrastructure to Meet the Challenges of Heating Electrification'

<p>The codes and data used in the followng paper</p> <p>''Khorramfar, R., Santoni-Calvin, M., Mallapragada, D., Amin, S., Botterud, A.,<br>Norfork L., (2025) Cost-effective Planning of Power-Gas Infrastructure to Meet the Challenges<br>of Heating Electrification, Cell Reports Sustainability</p> <p>&nbsp;</p> <p>Link (open source): https://www.cell.com/cell-reports-sustainability/fulltext/S2949-7906(25)00003-5</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Webinar: Potential and challenges of certification and development of new labels

<p>The webinar will address the vital topic of new certifications and labels for the renewable, and especially bio-based, economy.</p> <p>Certification and labelling play a crucial role in empowering consumers to make informed and sustainable purchasing decisions by impeding greenwashing and instead providing credible and reliable sustainability information. Furthermore, they facilitate the tracking and traceability of biological and other renewable feedstock throughout value chains, fostering transparency and accountability among stakeholders of the industry. However, there are a lot of challenges for current certification and labelling schemes, since they are often not sufficiently laid-out for bio-based and other renewable products and their value-chains.</p> <p>In this webinar, REDcert and T&Uuml;V AUSTRIA will show their sytems for sustainability certification and will present new or soon to be published label and certification schemes (LCS) that are relevant for the bio-based economy. The speakers will talk about challenges on the way to the final label, and which gaps will be closed.</p> <p>Join us on November the 5th and listen to the two presentations by REDcert and T&Uuml;V AUSTRIA. Together, we will explore their innovative new LCS that will support shaping a strong circular EU (bio)economy.</p> <p>Speakers:</p> <p>Phillipe Dewolfs (T&Uuml;V AUSTRIA): From standardisation to communication &ndash; the role of a certification body<br>Simon Schwarzwald (REDcert): Certification of sustainable products in the REDcert&sup2; scheme</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Model weights for a Weather4cast 2021 Challenge Stage 1 solution

<p>This repository contains the pre-trained model weights for the TensorFlow/Keras models used in the <a href="https://www.iarai.ac.at/weather4cast/2021-competition/challenge/">Weather4cast 2021 Challenge Stage 1</a> by the team &quot;antfugue&quot;. The model code can be found in <a href="https://github.com/jleinonen/weather4cast-stage1">https://github.com/jleinonen/weather4cast-stage1</a> along with instructions on where to extract the weights.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

WikiChurches – A Fine-Grained Dataset of Architectural Styles with Real-World Challenges

<p>WikiChurches is a dataset for architectural style classification, consisting of 9,485 images of church buildings. Both images and style labels were sourced from Wikipedia. The dataset can serve as a benchmark for various research fields, as it combines numerous real-world challenges: fine-grained distinctions between classes based on subtle visual features, a comparatively small sample size, a highly imbalanced class distribution, a high variance of viewpoints, and a hierarchical organization of labels, where only some images are labeled at the most precise level. In addition, we provide 631 bounding box annotations of characteristic visual features for 139 churches from four major categories. These annotations can, for example, be useful for research on fine-grained classification, where additional expert knowledge about distinctive object parts is often available.</p> <p>Please refer to the README.md file for information about the different files contained in this dataset.</p>

opencc-by-sa-4.0Aug 2021View details →
zenodo44/100

Data and Code from: On-farm land management strategies and production challenges in United States Organic Agricultural Systems.

<p>This repository contains data and code used in:</p> <p>Isaac Mpanga, Russel Trondstad, Jessica Guo, David LeBauer, and John Omololu, 2021. On-farm land management strategies and production challenges in United States Organic Agricultural Systems. Current Research in Environmental Sustainability.</p> <p>It provides USDA Surveys of Agricultural Production from 2008-2019 to investigate state and national trends by state in organic farm area, number, and sales, as well to evaluate national trends in on-farm land-use practices and challenges facing US organic production.</p> <p>It also includes code used to transform, visualize, and analyze the data, and derived data products - notably organic farm area and sales with values imputed to correct for redacted state level measures.</p>

openmit-licenseOct 2021View details →
zenodo44/100

Document Liveness Challenge (DLC-2021) - part 1 (or, cg)

<p>Dataset DLC-2021 consists of 1424 video clips captured in a wide range of real-world conditions and focused on ID document forensics tasks.&nbsp;Each clip was shot vertically and was at least 5 seconds long. Frames extracted at 10 frames per second and for the 50 first extracted frames document position is manually annotated.<br> The novelty of the dataset is that it contains shots from video with color laminated mock ID documents, color unlaminated copies, grayscale unlaminated copies, and screen recaptures of the documents.&nbsp;The proposed dataset complies with the GDPR because it contains images of synthetic IDs with generated owner photos and artificial personal information.</p> <p>Part 1 contains videos, frames and markup for &ldquo;original&rdquo; laminated documents from MIDV-2020 collection and unlaminated gray copies.&nbsp;<br> Part 2 contains videos, frames and markup for documents recaptured from device screen<br> Part 3 contains videos, frames and markup for unlaminated color copies.</p> <p><strong>Share and Cite</strong></p> <p><em>MDPI and ACS Style</em></p> <p>Polevoy, D.V.; Sigareva, I.V.; Ershova, D.M.; Arlazarov, V.V.; Nikolaev, D.P.; Ming, Z.; Luqman, M.M.; Burie, J.-C. Document Liveness Challenge Dataset (DLC-2021).&nbsp;<em>J. Imaging</em>&nbsp;<strong>2022</strong>,&nbsp;<em>8</em>, 181. https://doi.org/10.3390/jimaging8070181</p> <p><em>AMA Style</em></p> <p>Polevoy DV, Sigareva IV, Ershova DM, Arlazarov VV, Nikolaev DP, Ming Z, Luqman MM, Burie J-C. Document Liveness Challenge Dataset (DLC-2021).&nbsp;<em>Journal of Imaging</em>. 2022; 8(7):181. https://doi.org/10.3390/jimaging8070181</p> <p><em>Chicago/Turabian Style</em></p> <p>Polevoy, Dmitry V., Irina V. Sigareva, Daria M. Ershova, Vladimir V. Arlazarov, Dmitry P. Nikolaev, Zuheng Ming, Muhammad M. Luqman, and Jean-Christophe Burie. 2022. &quot;Document Liveness Challenge Dataset (DLC-2021)&quot;&nbsp;<em>Journal of Imaging</em>&nbsp;8, no. 7: 181.&nbsp; https://doi.org/10.3390/jimaging8070181</p>

opencc-by-sa-2.5Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record