Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,863

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,863 results for “Challenge”

Learn how ShareScore rates datasets ↗
zenodo44/100

The HumBug Challenge: ComParE 2022

<p><strong>A large-scale multi-species dataset of acoustic recordings</strong></p> <p>Dataset compatible with two papers:</p> <ul> <li>The&nbsp;<strong>Computational Paralinguistics ChallengE (ComParE): Mosquito Event Detection Task</strong><strong>&nbsp;</strong><a href="https://github.com/EIHW/ComParE2022/tree/MOS-C">https://github.com/EIHW/ComParE2022/tree/MOS-C</a></li> <li>An update to:&nbsp;<em>HumBugDB: a large-scale acoustic mosquito dataset:</em> <ul> <li><a href="https://arxiv.org/abs/2110.07607">NeurIPS 2021 Paper</a></li> <li><a href="https://github.com/HumBug-Mosquito/HumBugDB">https://github.com/HumBug-Mosquito/HumBugDB</a>.</li> </ul> </li> </ul> <p>A large-scale multi-species dataset containing recordings of mosquitoes collected from multiple locations globally, as well as via different collection methods.&nbsp;In total, we present&nbsp;20 hours&nbsp;of labelled mosquito data with 15 hours&nbsp;of corresponding background noise, recorded at the sites of 8 experiments.&nbsp;&nbsp;Of these, 64,843 seconds contain species metadata, consisting of 36 species (or species complexes).</p> <p>This repository contains:</p> <ul> <li>Audio files to be extracted into <em>audio/data/</em><em>train</em>&nbsp;and&nbsp;<em>audio/data/</em><em>dev/{a/b} respectively</em></li> <li>Metadata in <em>csv</em> format:&nbsp;<a href="https://zenodo.org/record/6478589/files/humbugdb_zenodo_0_0_2.csv?download=1">neurips_2021_zenodo_0_0_2.csv</a></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Online survey of needs and challenges of innovation ecosystems and intermediaries for taking up activity in the EU space sector

<p>The present dataset was generated as part of the &quot;Needs and challenges of innovation ecosystems and intermediaries for taking up activity in the EU space sector&quot; of the H2020 <a href="http://innorbit.eu">InnORBIT project</a>.</p> <p>The aim of this study was to identify and explore the available and missing skills of innovation intermediaries to provide business support services to innovators within their local ecosystems to develop commercial activity in space. The assessment of skills was based on a baseline framework encompassing a wide array of skills and competencies innovation intermediaries are supposed to possess in order to provide effective business support services to space innovators. The skills of the baseline framework are&nbsp;grouped into five broad categories: (i) space industry knowledge, (ii) business assessment knowledge, (iii) business support skills, (iv) organisational and digital skills and (v) soft skills. The baseline framework was originally developed by the InnORBIT consortium through research in related works of EU&nbsp;funded projects and publications and validated through a series of 15 interviews with top-level executives of organisations across the CEE and SEE area,&nbsp;belonging to the two target groups of the study (i.e., innovation intermediaries and innovators). An online survey was deployed from May 26th to June 18th using the EU Survey tool, to innovation intermediaries and innovators across the EU and CEE/SEE countries in particular.&nbsp;Two online questionnaires were developed building on the baseline skills framework - the first intended for innovation intermediaries asking them to perform a self-assessment of their skills in terms of providing business support services and the second targeting innovators, asking them to state their perception on how innovation intermediaries they have worked with, perform in each of the skills.</p> <p>The dataset contains four files:</p> <p>1. Zip file including the transcripts from 6&nbsp;interviews with space innovators in Eastern Europe for the evaluation of the baseline framework of skills.</p> <p>2. Zip file including the transcripts from 9 interviews with innovation intermediaries in Eastern Europe for the evaluation of the baseline framework of skills.</p> <p>3. a pdf file of the digital&nbsp;questionnaires developed in EU Survey deployed to innovation intermediaries and innovators in the region</p> <p>4. An excel file with&nbsp;104 valid responses collected from the online survey (56 innovation intermediaries and 48 innovators) across 16 EU countries / 21 countries total.</p> <p>The dataset contains only non-sensitive anonymised information and is in full compliance with the GDPR provisions. Any information&nbsp;leading to the identification of participants in activities (interviews, survey) is either modified or omitted and deonted with brackets.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

COST-ARKWORK Societal Challenges Survey Dataset

<p>A survey was relating to the societal role of archaeology and the impact of societal challenges on archaeological practices is based on views collected from the members of COST Action Archaeological Practices and Knowledge Work in the Digital Environment (<a href="http://www.arkwork.eu/">www.arkwork.eu</a>). The purpose of this survey is to collect views from COST-ARKWORK members on the relation of contemporary societal challenges and archaeological practices. The survey was administered online using Survey &amp; Report tool (survey form, Fig. 1) hosted by&nbsp;Swedish University Computer Network (Sunet). The survey was open from May 8, 2020 until June1, 2020.&nbsp;</p> <p>&nbsp;</p> <p>The participants of the network consist of over 200 experts of archaeological practices from a broad range of disciplinary backgrounds from archaeology to information science, museum studies, computer science and business studies, including researchers and practitioners from 30 European countries. 50 members of the network participated in the survey. The views of the Action participants were collected using an online survey with two open ended questions: 1. &ldquo;From your perspective, what is the role and value of archaeology in helping to solve contemporary societal problems?&rdquo;; and, 2. &ldquo;What societal changes and challenges are likely to affect archaeology and archaeological practices the most during the next five years?&rdquo;. In addition, the respondents were asked to indicate whether they identified themselves as being an archaeologist or not. 64% (32/50) of the experts identified themselves as archaeologists while the rest described themselves as non-archaeologists. The experts shared their opinions as individuals, not as representants of specific institutions or countries.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Neuromotor dynamics of human locomotion in challenging settings

<p><strong>Background</strong></p> <p>Is the control of movement less stable when we walk or run in challenging settings? Intuitively, one might answer that it is, given that challenging locomotion externally (e.g. rough terrain) or internally (e.g. age-related impairments) makes our movements more unstable. Here, we investigated how young and old humans synergistically activate muscles during locomotion when different perturbation levels are introduced. Of these control signals, called muscle synergies, we analyzed the stability over time and the complexity (or irregularity). Surprisingly, we found that perturbations force the central nervous system to produce muscle activation patterns that are less unstable and less complex. These outcomes show that robust locomotion in challenging settings is achieved by producing less complex control signals which are more stable over time, whereas easier tasks allow for more unstable and irregular control.</p> <p><strong>How to use the data set</strong></p> <p>This supplementary data set contains: a) the metadata with anonymized participant information, b) the raw electromyographic (EMG) data acquired during locomotion, c) the touchdown and lift-off timings of the recorded limb, d) the filtered and time-normalized EMG, e) the muscle synergies extracted via non-negative matrix factorization and f) the code written in R (R Found. for Stat. Comp.) to process the data, including the scripts to calculate the short-term Maximum Lyapunov Exponents (sMLE) and Higuchi&#39;s fractal dimension (HFD) of motor primitives. In total, 476 trials from 86 participants are included in the supplementary data set.</p> <p>The file &ldquo;participant_data.dat&rdquo; is available in ASCII and RData (R Found. for Stat. Comp.) format and contains:</p> <ul> <li>Code: the participant&rsquo;s code</li> <li>Experiment: the experimental setup in which the participant was involved (E1 = walking and running, overground and treadmill; E2 = walking and running, even- and uneven-surface; E3 = unperturbed and perturbed walking, young and old)</li> <li>Group: the group to which the participant was assigned (see methods for the details)</li> <li>Sex: the participant&rsquo;s sex (M or F)</li> <li>Speed: the speed at which the recordings were conducted in [m/s] (two values separated by a comma mean that recordings were done at two different speeds, i.e. walking and running)</li> <li>Age: the participant&rsquo;s age in years (participants were considered old if older than 65 years, but younger than 80)</li> <li>Height: the participant&rsquo;s height in [cm]</li> <li>Mass: the participant&rsquo;s body mass in [kg].</li> </ul> <p>The &quot;RAW_DATA.RData&quot;&nbsp;R list consists of elements of S3 class &quot;EMG&quot;, each of which is a human locomotion trial containing cycle segmentation timings and raw electromyographic (EMG) data from 13 muscles of the right-side leg. Cycle times are structured as data frames containing two columns that&nbsp;correspond to touchdown (first column) and lift-off (second column).&nbsp;Raw EMG data sets are also structured as data frames with one row for each recorded data point&nbsp;and 14 columns. The first column contains the incremental time in seconds. The remaining 13 columns contain the raw EMG data, named with the following muscle abbreviations:&nbsp;ME = gluteus medius, MA = gluteus maximus, FL = tensor fasci&aelig; lat&aelig;, RF = rectus femoris, VM = vastus medialis, VL = vastus lateralis, ST = semitendinosus, BF = biceps femoris, TA = tibialis anterior, PL = peroneus longus, GM = gastrocnemius medialis, GL = gastrocnemius lateralis, SO = soleus. Please note that the running overground trials of participants P0001 and P0009 and the second uneven-surface running trial of participant P0048 consist of 22, 27 and 23 cycles, respectively. All the other trials&nbsp;consist of 30 gait cycles. Trials are named like &ldquo;P0053_OW_02&rdquo;, where the characters&nbsp;&ldquo;P0053&rdquo; indicate the participant number (in this example the 53rd), the characters &ldquo;OW&rdquo; indicate the locomotion type (E1: OW=overground walking, OR=overground running, TW=treadmill walking, TR=treadmill running; E2: EW=even-surface walking, ER=even-surface running, UW=uneven-surface walking, UR=uneven-surface running; E3: NW=normal walking, PW=perturbed walking), and the numbers &ldquo;02&rdquo; indicate the trial number (in this case the 2nd). The 10 trials per participant recorded for each overground session (i.e. 10 for walking and 10 for running) were concatenated into one. The filtered and time-normalized EMG data are named, following the same rules, like &ldquo;FILT_EMG_P0053_OG_02&rdquo;.</p> <p><strong>Old versions not compatible with the R package <a href="https://CRAN.R-project.org/package=musclesyneRgies">musclesyneRgies</a></strong></p> <p>The files containing the gait cycle breakdown are available in RData (R Found. for Stat. Comp.) format, in the file named &ldquo;CYCLE_TIMES.RData&rdquo;. The files are structured as data frames with 30 rows (one for each gait cycle) and two columns. The first column contains the touchdown incremental times in seconds. The second column contains the duration of each stance phase in seconds. Each trial is saved as an element of a single R list. Trials are named like &ldquo;CYCLE_TIMES_P0020,&rdquo; where the characters &ldquo;CYCLE_TIMES&rdquo; indicate that the trial contains the gait cycle breakdown times and the characters &ldquo;P0020&rdquo; indicate the participant number (in this example the 20th). Please note that the overground trials of participants P0001 and P0009 and the second uneven-surface running trial of participant P0048 only contain 22, 27 and 23 cycles, respectively.</p> <p>The files containing the raw, filtered and the normalized EMG data are available in RData (R Found. for Stat. Comp.) format, in the files named &ldquo;RAW_EMG.RData&rdquo; and &ldquo;FILT_EMG.RData&rdquo;. The raw EMG files are structured as data frames with 30000 rows (one for each recorded data point) and 14 columns. The first column contains the incremental time in seconds. The remaining thirteen columns contain the raw EMG data, named with muscle abbreviations that follow those reported in the methods section. Each trial is saved as an element of a single R list. Trials are named like &ldquo;RAW_EMG_P0053_OW_02&rdquo;, where the characters &ldquo;RAW_EMG&rdquo; indicate that the trial contains raw emg data, the characters &ldquo;P0053&rdquo; indicate the participant number (in this example the 53rd), the characters &ldquo;OW&rdquo; indicate the locomotion type (E1: OW=overground walking, OR=overground running, TW=treadmill walking, TR=treadmill running; E2: EW=even-surface walking, ER=even-surface running, UW=uneven-surface walking, UR=uneven-surface running; E3: NW=normal walking, PW=perturbed walking), and the numbers &ldquo;02&rdquo; indicate the trial number (in this case the 2nd). The 10 trials per participant recorded for each overground session (i.e. 10 for walking and 10 for running) were concatenated into one. The filtered and time-normalized EMG data is named, following the same rules, like &ldquo;FILT_EMG_P0053_OG_02&rdquo;.</p> <p>The files containing the muscle synergies extracted from the filtered and normalized EMG data are available in RData (R Found. for Stat. Comp.) format, in the files named &ldquo;SYNS_H.RData&rdquo; and &ldquo;SYNS_W.RData&rdquo;. The muscle synergies files are divided in motor primitives and motor modules and are presented as direct output of the factorization and not in any functional order. Motor primitives are data frames with 6000 rows and a number of columns equal to the number of synergies (which might differ from trial to trial) plus one. The rows contain the time-dependent coefficients (motor primitives), one column for each synergy plus the time points (columns are named e.g. &ldquo;Time, Syn1, Syn2, Syn3&rdquo;, where &ldquo;Syn&rdquo; is the abbreviation for &ldquo;synergy&rdquo;). Each gait cycle contains 200 data points, 100 for the stance and 100 for the swing phase which, multiplied by the 30 recorded cycles, result in 6000 data points distributed in as many rows. This output is transposed as compared to the one discussed above to improve user readability. Each set of motor primitives is saved as an element of a single R list. Trials are named like &ldquo;SYNS_H_P0012_PW_02&rdquo;, where the characters &ldquo;SYNS_H&rdquo; indicate that the trial contains motor primitive data, the characters &ldquo;P0012&rdquo; indicate the participant number (in this example the 12th), ), the characters &ldquo;PW&rdquo; indicate the locomotion type (see above), and the numbers &ldquo;02&rdquo; indicate the trial number (in this case the 2nd). Motor modules are data frames with 13 rows (number of recorded muscles) and a number of columns equal to the number of synergies (which might differ from trial to trial). The rows, named with muscle abbreviations that follow those reported in the methods section, contain the time-independent coefficients (motor modules), one for each synergy and for each muscle. Each set of motor modules relative to one synergy is saved as an element of a single R list. Trials are named like &ldquo;SYNS_W_P0082_PW_02&rdquo;, where the characters &ldquo;SYNS_W&rdquo; indicate that the trial contains motor module data, the characters &ldquo;P0082&rdquo; indicate the participant number (in this example the 82nd) ), the characters &ldquo;PW&rdquo; indicate the locomotion type (see above), and the numbers &ldquo;02&rdquo; indicate the trial number (in this case the 2nd). Given the nature of the NMF algorithm for the extraction of muscle synergies, the supplementary data set might show non-significant differences as compared to the one used for obtaining the results of this paper.</p> <p>The files containing the sMLE calculated from motor primitives are available in RData (R Found. for Stat. Comp.) format, in the file named &ldquo;sMLE.RData&rdquo;. sMLE results are presented in a list of lists containing, for each trial, 1) the divergences, 2) the sMLE, and 3) the value of the R<sup>2</sup> between the divergence curve and its linear interpolation made using the specified amount of points. The divergences are presented as a one-dimensional vector. sMLE are one number like the R<sup>2</sup> value. Trials are named like &ldquo;MLE_P0081_EW_01&rdquo;, where the characters &ldquo;sMLE&rdquo; indicate that the trial containss sMLE data, the characters &ldquo;P0081&rdquo; indicate the participant number (in this example the 81st) ), the characters &ldquo;EW&rdquo; indicate the locomotion type (see above), and the numbers &ldquo;01&rdquo; indicate the trial number (in this case the 1st).</p> <p>The files containing the HFD calculated from motor primitives are available in RData (R Found. for Stat. Comp.) format, in the file named &ldquo;HFD.RData&rdquo;. HFD results are presented in a list of lists containing, for each trial, 1) the HFD, and 2) the interval time k used for the calculations. HFDs are presented as one number, as are the interval times k. Trials are named like &ldquo;HFD_P0048_TR_01&rdquo;, where the characters &ldquo;HFD&rdquo; indicate that the trial contains HFD data, the characters &ldquo;P0048&rdquo; indicate the participant number (in this example the 48th), the characters &ldquo;TR&rdquo; indicate the locomotion type (see above), and the numbers &ldquo;01&rdquo; indicate the trial number (in this case the 1st).</p> <p>All the code used for the preprocessing of EMG data, the extraction of muscle synergies, the calculation of sMLE and HFD is available in R (R Found. for Stat. Comp.) format. Explanatory comments are profusely present throughout the scripts (&ldquo;SYNS.R&rdquo;, which is the script to extract synergies, &ldquo;fun_NMF.R&rdquo;, which contains the NMF function, &ldquo;sMLE.R&rdquo;, which is the script to calculate the sMLE of motor primitives, &ldquo;HFD.R&rdquo;, which is the script to calculate the HFD of motor primitives, &ldquo;fun_sMLE.R&rdquo;, which contains the sMLE function and &ldquo;fun_HFD.R&rdquo;, which contains the HFD function).</p>

opencc-by-4.0Sep 2019View details →
zenodo44/100

Numerical data analysed to produce Figure 3a of Nature Climate Change submission "Five challenges for subseasonal to decadal prediction research " by Merryfield et al.

<p>NetCDF4-formatted files containing daily sea ice concentration data from Environment and Climate Change Canada&#39;s CanSIPSv2&nbsp;seasonal forecasting system described in Lin et al. (2020)&nbsp;https://doi.org/10.1175/WAF-D-19-0259.1&nbsp;</p> <ul> <li>2 models, CanCM4i and GEM-NEMO</li> <li>10 ensemble members for&nbsp;each model, each in separate files as indicated by suffixes _1 to _10</li> <li>initialized May 1, 1980 to 2021</li> <li>840 files total (42 predicted years x 10 ensemble members x 2 models)</li> <li>model outputs interpolated to common 1-degree grid</li> </ul> <p>The calibrated probabilistic forecast map shown in Figure 3a is based on&nbsp;the&nbsp;nonhomogeneous censored Gaussian regression (NCGR) method described in Dirkson et al, (2021)&nbsp;https://doi.org/10.1175/WAF-D-20-0066.1 and produced using scripts available at&nbsp;https://github.com/adirkson/sea-ice-timing&nbsp;</p> <p>The procedure&nbsp;uses as inputs</p> <ul> <li>freeze-up dates calculated from the provided model outputs as described in Sigmond et al. (2016)&nbsp;https://doi.org/10.1002/2016GL071396</li> <li> <p>NOAA/NSIDC Climate Data Record of Passive Microwave Sea Ice Concentration, Version 3: https://nsidc.org/data/G02202/versions/3</p> </li> </ul>

opencc-by-4.0Jul 2022View details →
zenodo44/100

LSSTC AGN Data Challenge 2021

<p>This repository hosts the dataset used in the LSSTC AGN Data Challenge (DC) 2021 (PI: Gordon&nbsp;Richards). More information about the data challenge can be found in the DC GitHub repository @&nbsp;https://github.com/RichardsGroup/AGN_DataChallenge.&nbsp;</p> <p><strong>Dataset Versions:&nbsp;</strong></p> <p><strong>1.0:&nbsp;</strong>The initial dataset used in the DC, as well as the blinded dataset (ObjectTable_Blinded.parquet) that was&nbsp;used to evaluate submissions. Note that the image cutouts are not included here due to the large size, but the script used to generate those cutouts using SDSS archive services&nbsp;is included&nbsp;in the DC&nbsp;GitHub repository.</p> <p><strong>1.1:&nbsp;</strong>The same dataset as in v1.0 but with the following updates:</p> <ul> <li>Uncovered the true coordinates of each source in the dataset</li> <li>Added E(B-V) for every source using the SFD1998 dust map</li> <li>Added spectrum source information (i.e., SDSS fiberid, plate, mjd) if available.&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Caveat:</strong>&nbsp;</p> <ul> <li>The optical (grizY) and NIR photometry of sources in the XMM-LSS field is a product&nbsp;of the HSC/VISTA pixel-level joint processing initiative led by Raphael Shirley and Manda Banerji. Thus, it is&nbsp;an early prototype dataset and is still subject to testing and characterization.</li> </ul> <p>&nbsp;</p> <p><strong>Citation:</strong></p> <p>The DC dataset released here is a compilation of data from various sources. If you find the DC dataset useful for your research and would like to acknowledge it, please also reference the original sources of the data. Below is a list of publications that you should consider citing.&nbsp;</p> <p><em>X-ray in XMM-LSS (XMM-SERVS):&nbsp;</em>2018MNRAS.478.2132C</p> <p><em>UV Photometry (GALEX):&nbsp;</em>2017ApJS..230...24B</p> <p><em>Optical Photometry (in the object/source tables):&nbsp;</em></p> <ul> <li>DES: 2021ApJS..255...20A</li> <li>SDSS Stripe 82 Coadd:&nbsp;2014ApJ...794..120A</li> <li>HSC DR2:&nbsp;2019PASJ...71..114A</li> </ul> <p><em>Optical Light Curves (in the ForcedSource table):</em></p> <ul> <li>SDSS DR7:&nbsp;2009ApJS..182..543A</li> <li>SDSS II Supernova Survey:&nbsp;2008AJ....135..338F</li> </ul> <p><em>Astrometry (i.e., parallax, proper motion):</em></p> <ul> <li>Gaia EDR3:&nbsp;2021A&amp;A...649A...1G</li> <li>NOIRLab Source Catalog DR2:&nbsp;2021AJ....161..192N</li> </ul> <p><em>NIR in XMM-LSS (VISTA/VIDEO):&nbsp;</em>2013MNRAS.428.1281J</p> <p><em>NIR in Stripe 82 (UKIDSS):</em></p> <ul> <li> <p>2006MNRAS.367..454H</p> </li> <li> <p>2007MNRAS.379.1599L</p> </li> <li> <p>2008MNRAS.384..637H</p> </li> <li> <p>2009MNRAS.394..675H</p> </li> </ul> <p><em>Optical u-band in XMM-LSS (CFHTLS):</em>&nbsp;2012yCat.2317....0H</p> <p><em>MIR in XMM-LSS </em>(Spitzer&nbsp;DeepDrill):&nbsp;2021MNRAS.501..892L</p> <p><em>MIR in Stripe 82 </em>(SpIES):&nbsp;2016ApJS..225....1T</p> <p><em>FIR</em> (Hershel/HELP):&nbsp;2019MNRAS.490..634S</p> <p><em>Radio</em> (FIRST):&nbsp;1994ASPC...61..165B</p> <p><em>HighZ QSOs:&nbsp;</em></p> <ul> <li>2016ApJ...819...24W</li> <li>2016ApJ...829...33Y</li> </ul> <p><em>SDSS Spectroscopy:</em></p> <ul> <li>SDSS DR16:&nbsp;2020ApJS..249....3A</li> <li>SDSS DR16 Quasar Catalog:&nbsp;2020ApJS..250....8L</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Exoplanet imaging data challenge, phase 2

<p><strong>Datasets for the second phase of the Exoplanet Imaging Data Challenge</strong> (<a href="https://exoplanet-imaging-challenge.github.io/">https://exoplanet-imaging-challenge.github.io</a>).&nbsp;</p> <p>The second phase of the Exoplanet Imaging Data Challenge is focused on the characterisation of exoplanet signals in high-contrast imaging data. The participants must perform two tasks: provide (1) astrometry of the detected signals, and (2) spectrophotometry of the detected signals.</p> <p>For this phase, we therefore provide with <strong>8</strong> high-contrast data sets, taken with two integral field spectrographs: SPHERE-IFS installed at the Very Large Telescope (VLT, Chile) and GPI installed at the Gemini-South telescope (Chile). Each data set consists of the following files (in <em>.fits</em> format):</p> <p>(a) <em>image_cube_instxx.fits</em>: the coronagraphic multispectral image cube, acquired in pupil-stabilized mode;<br> (b) <em>parallactic_angles_instxx.fits</em>: the corresponding parallactic angle and airmass variation during the observation sequence;<br> (c)&nbsp;<em>wavelength_vect_instxx.fits</em>: the corresponding wavelength vector for each spectral channel;<br> (d) <em>psf_cube_instxx.fits</em>: the non-coronagraphic (point spread function) multispectral image of the target star;<br> (e) <em>first_guess_astrometry_instxx.fits</em>: a first guess position w.r.t the star of the (2 or 3) injected planetary signals.</p> <p>Each data set has 2 to 3 injected planetary signals in various locations.<br> Each data set has very different observing conditions, from very good to very bad.</p> <p><em>Additional information and ressources can be found on the <a href="https://exoplanet-imaging-challenge.github.io/">website</a> and dedicated <a href="https://github.com/exoplanet-imaging-challenge/phase2">Github repository.</a></em></p> <p><em>The results of the data challenge must be submitted directly by participants on the <a href="https://eval.ai/web/challenges/challenge-page/1717/overview">EvalAI platform</a>.</em></p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

LISA Data Challenge Sangria (LDC2a)

<p><strong>Sangria</strong> includes two main datasets: each contains Gaussian instrumental noise and simulated waveforms from 30 million Galactic white dwarf binaries, from 17 verification Galactic binaries, and from merging massive black-hole binaries with parameters derived from an astrophysical model. The first dataset includes the full specification used to generate it: source parameters, a description of instrumental noise with the corresponding power spectral density, LISA&#39;s orbit, etc. We also release noiseless data for each type of source, for waveform validation purposes. The second dataset is <em>blinded</em>: the level of istrumental noise and number of sources of each type are not disclosed (except for the known parameters of the verification binaries).</p> <p>See <a href="https://lisa-ldc.lal.in2p3.fr/challenge2a">LDC website</a> for more details.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Challenges of high-fidelity air quality modeling in urban environments - PALM sensitivity study during stable conditions (TURBAN)

<h3>Introduction</h3> <p>This dataset contains the PALM model inputs and the source code used to create the simulations for Prague-Legerova scenarios performed in the scope of the&nbsp;<strong>TURBAN</strong> project (<a href="https://www.project-turban.eu/">https://www.project-turban.eu/</a>). Detailed description of the simulations is provided in the referencing scientific paper.</p> <h3>List of simulations</h3> <table> <tbody> <tr> <td><strong>Scenario name</strong></td> <td><strong>Days simulated</strong></td> <td><strong>IBC</strong></td> <td><strong>Configuration changes</strong></td> </tr> <tr> <td>legerovas_s6_sens_base</td> <td>13&ndash;15 February 2023</td> <td>ICON</td> <td>-</td> </tr> <tr> <td>legerovas_s6_sens_dtmax</td> <td>13 February 2023</td> <td>ICON</td> <td>dt_max=0.2</td> </tr> <tr> <td>legerovas_s6_sens_heat</td> <td>13 February 2023</td> <td>ICON</td> <td>car anthropogenic heat (custom code)</td> </tr> <tr> <td>legerovas_s6_sens_sgs</td> <td>13 February 2023</td> <td>ICON</td> <td>e_min=0.02</td> </tr> <tr> <td>legerovas_s6_sens_stg</td> <td>13 February 2023</td> <td>ICON</td> <td>STG_PROFILES added</td> </tr> <tr> <td>legerovas_s6_sens_alad</td> <td>13&ndash;15 February 2023</td> <td>ALADIN</td> <td>-</td> </tr> <tr> <td>legerovas_s6_sens_alad_heat</td> <td>13 February 2023</td> <td>ALADIN</td> <td>car anthropogenic heat (custom code)</td> </tr> <tr> <td>legerovas_s6_sens_alad_sgs</td> <td>13 February 2023</td> <td>ALADIN</td> <td>e_min=0.02</td> </tr> <tr> <td>legerovas_s6_sens_alad_stg</td> <td>13 February 2023</td> <td>ALADIN</td> <td>STG_PROFILES added</td> </tr> <tr> <td>legerovas_s6_sens_wrf</td> <td>13&ndash;15 February 2023</td> <td>WRF</td> <td>-</td> </tr> <tr> <td>legerovas_s6_sens_wrf_heat</td> <td>13 February 2023</td> <td>WRF</td> <td>car anthropogenic heat (custom code)</td> </tr> <tr> <td>legerovas_s6_sens_wrf_sgs</td> <td>13 February 2023</td> <td>WRF</td> <td>e_min=0.02</td> </tr> <tr> <td>legerovas_s6_sens_wrf_stg</td> <td>13 February 2023</td> <td>WRF</td> <td>STG_PROFILES added</td> </tr> </tbody> </table> <h3>Directory structure</h3> <p>The directory inputs contains the model inputs and it is further divided into these subdirectories:</p> <p>- inputs/common: The PALM static driver and the emission drivers for the parent and child domains. These files are common to all simulations</p> <p>- inputs/dynamic/*: These directories contain the dynamic drivers for the parent and child domanis, which contain the initial and boundary conditions (IBC) as well as external radiation data. The three subdirectories aladin, icon and wrf contain IBCs created from the respective mesoscale model outputs.&nbsp;</p> <p>- inputs/legerovas_s6_sens_*: These directories contain the PALM model configuration (p3d) for both domains for each simulation.</p> <p>- inputs/build_config: The included .palm.iofiles configuration file ensures that the files STG_PROFILES are correctly copied from the input directory.</p> <p>The directory palm_sources contains the exact model source used for the simulations. It is derived from the PALM model release 23.04 with additional bugfixes. There are two source archives:</p> <p>- heat.tar.gz: PALM source further modified to include anthropogenic heat from cars, used for the simulations legerovas_s6_sens_*_heat</p> <p>- standard.tar.gz: PALM source used for all other included simulations.</p> <h3>Reproducing the simulations</h3> <p>In order to reproduce the simulations, unpack the respective source code archive and follow the standard installation, configuration and build procedures described in the README.md file within the archive and on the PALM model website http://www.palm-model.org/. Then copy the input files for the respective simulation in the JOBS directory. The common files and the dynamic driver files need to be renamed so that they match the prefix given by the name of the simulation, as is described in the PALM model documentation.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

The PANORAMA Challenge: Public Training and Development Dataset (3)

<p>This dataset represents the <strong><a href="https://panorama.grand-challenge.org/" target="_blank" rel="noopener">PANORAMA</a>: Public Training and Development Dataset</strong>. It contains 2238 anonymized contrast-enhanced CT (CECT) scans acquired at two centers (Radboud University Medical Center, University Medical Center Groningen) based in The Netherlands. Additionally, it contains 194 cases from the&nbsp;<strong><a href="http://medicaldecathlon.com/" target="_blank" rel="noopener">Medical Segmentation Decathlon</a>&nbsp;</strong>dataset and 80 cases from<strong>&nbsp;<a href="https://www.cancerimagingarchive.net/collection/pancreas-ct/" target="_blank" rel="noopener">National Institutes of Health</a></strong>. For all updates/fixes regarding this dataset, please join the challenge and check out our&nbsp;<a href="https://grand-challenge.org/forums/forum/panorama-pancreatic-cancer-diagnosis-radiologists-meet-ai-711/topic/public-training-and-development-dataset-updates-and-fixes-2213/" target="_blank" rel="noopener">dedicated forum post</a> on this topic. The corresponding labels of the PANORAMA dataset can be found <a href="https://github.com/DIAGNijmegen/panorama_labels">here</a>.&nbsp;</p> <p>The PANORAMA challenge is an all-new grand challenge that aims to validate the diagnostic performance of artificial intelligence and radiologists at pancreatic ductal adenocarcinoma (PDAC) detection/diagnosis in CECT, with histopathology and follow-up (&ge; 3 years) as the reference standard, in a retrospective setting in the hidden testing dataset. The study hypothesizes that state-of-the-art AI algorithms are non-inferior to radiologists reading CECT.</p> <p>Key aspects of the PANORAMA study design have been established in conjunction with an international scientific advisory board of 13 experts in AI and pancreas radiology as well as a patient representative &mdash;to unify and standardize present-day guidelines, and to ensure meaningful validation of pancreas AI towards clinical translation (<strong><a href="https://www.sciencedirect.com/science/article/pii/S2405456921001607">Reinke et al., 2021</a></strong>).</p> <p><em>This PANORAMA dataset contains: batch&nbsp;<strong>3</strong> <strong>out of 4</strong></em></p>

opencc-by-nc-4.0Apr 2024View details →
zenodo44/100

The PANORAMA Challenge: Public Training and Development Dataset (4)

<p>This dataset represents the <strong><a href="https://panorama.grand-challenge.org/" target="_blank" rel="noopener">PANORAMA</a>: Public Training and Development Dataset</strong>. It contains 2238 anonymized contrast-enhanced CT (CECT) scans acquired at two centers (Radboud University Medical Center, University Medical Center Groningen) based in The Netherlands. Additionally, it contains 194 cases from the&nbsp;<strong><a href="http://medicaldecathlon.com/" target="_blank" rel="noopener">Medical Segmentation Decathlon</a>&nbsp;</strong>dataset and 80 cases from<strong>&nbsp;<a href="https://www.cancerimagingarchive.net/collection/pancreas-ct/" target="_blank" rel="noopener">National Institutes of Health</a></strong>. For all updates/fixes regarding this dataset, please join the challenge and check out our&nbsp;<a href="https://grand-challenge.org/forums/forum/panorama-pancreatic-cancer-diagnosis-radiologists-meet-ai-711/topic/public-training-and-development-dataset-updates-and-fixes-2213/" target="_blank" rel="noopener">dedicated forum post</a> on this topic. The corresponding labels of the PANORAMA dataset can be found <a href="https://github.com/DIAGNijmegen/panorama_labels">here</a>.&nbsp;</p> <p>The PANORAMA challenge is an all-new grand challenge that aims to validate the diagnostic performance of artificial intelligence and radiologists at pancreatic ductal adenocarcinoma (PDAC) detection/diagnosis in CECT, with histopathology and follow-up (&ge; 3 years) as the reference standard, in a retrospective setting in the hidden testing dataset. The study hypothesizes that state-of-the-art AI algorithms are non-inferior to radiologists reading CECT.</p> <p>Key aspects of the PANORAMA study design have been established in conjunction with an international scientific advisory board of 13 experts in AI and pancreas radiology as well as a patient representative &mdash;to unify and standardize present-day guidelines, and to ensure meaningful validation of pancreas AI towards clinical translation (<strong><a href="https://www.sciencedirect.com/science/article/pii/S2405456921001607">Reinke et al., 2021</a></strong>).</p> <p><em>This PANORAMA dataset contains: batch&nbsp;<strong>4</strong> <strong>out of 4</strong></em></p>

opencc-by-nc-4.0Apr 2024View details →
zenodo44/100

Trackerless 3D Freehand Ultrasound Reconstruction Challenge 2024 - Train Dataset (Part 2)

<blockquote> <p><strong>This Challenge will be an open-ended challenge, and we welcome your submission. Please register your team via this ⁠<a title="https://forms.office.com/e/dPg47ktV7M" href="https://forms.office.com/e/dPg47ktV7M" target="_blank" rel="noopener">form</a>. You can submit the algorithm via this <a title="https://forms.office.com/e/QChhNkLYiu" href="https://forms.office.com/e/QChhNkLYiu" target="_blank" rel="noopener noreferrer">form</a> for TUS-REC2024 Challenge, and we will test your submitted docker on the test set.</strong></p> <p><strong>We are organising TUS-REC2025 at MICCAI2025. More information is available on the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/" target="_blank" rel="noopener">TUS-REC2025 challenge website</a> and <a href="https://github.com/QiLi111/TUS-REC2025-Challenge_baseline" target="_blank" rel="noopener">Baseline code repo</a>.</strong></p> </blockquote> <p><strong>This is the second part of the Challenge dataset.&nbsp;<a href="../doi/10.5281/zenodo.11178509" target="_blank" rel="noopener">Link</a> to first part; <a href="../doi/10.5281/zenodo.11355500" target="_blank" rel="noopener">Link</a> to third part. <a href="../doi/10.5281/zenodo.12979481" target="_blank" rel="noopener">Link</a> to validation dataset.</strong></p> <p>Acquisition devices and config: The 2D US images were acquired using an Ultrasonix machine (BK, Europe) with a curvilinear probe (4DC7-3/40). The associated position information of each frame was recorded by an optical tracker (NDI Polaris Vicra, Northern Digital Inc., Canada). The acquired US frames were recorded at 20 fps, with an image size of 480&times;640, without speckle reduction. The frequency was set at 6MHz with a dynamic range of 83 dB, an overall gain of 48% and a depth of 9 cm.&nbsp;</p> <div> <p>Scanning protocol: Both left and right forearms of volunteers were scanned. For each forearm, the US probe moves in three different trajectories (straight line shape, "C" shape, and "S" shape), in a distal-to-proximal direction followed by a proximal-to-distal direction, with the US plane perpendicular of and parallel to the scanning direction. The train dataset contains 1200 scans in total, 24 scans associated with each subject.</p> <div> <div> <p>For detailed information please refer to the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/TUS-REC2024/" target="_blank" rel="noopener">Challenge website</a>. Baseline code is also provided, which can be found at this <a href="https://github.com/QiLi111/tus-rec-challenge_baseline" target="_blank" rel="noopener">repo</a>.</p> <p>Dataset structure:&nbsp;</p> </div> <div> <ul> <li> <p>The dataset contains 50 folders (one subject per folder), each with 24 scans. Each .h5 file corresponds to one scan, storing image and transformation of each frame within this scan. Key-value pairs in each .h5 file are explained below.</p> <ul> <li> <p>&ldquo;frames&rdquo;&nbsp; - All frames in the scan; with a shape of [N,H,W], where N refers to the number of frames in the scan, H and W denote the height and width of a frame.&nbsp;</p> </li> <li> <p>&ldquo;tforms&rdquo; - All transformations in the scan; with a shape of [N,4,4], where N is the number of frames in the scan, and the transformation matrix denotes the transformation from tracker tool space to camera space.&nbsp;</p> </li> <li> <p>Notations in the name of each .h5 file: &ldquo;RH&rdquo;: right arm; &ldquo;LH&rdquo;: left arm; &ldquo;Per&rdquo;: perpendicular; &ldquo;Par&rdquo;: parallel; &ldquo;L&rdquo;: straight line shape; &ldquo;C&rdquo;: C shape; &ldquo;S&rdquo;: S shape; &ldquo;DtP&rdquo;: distal-to-proximal direction; &ldquo;PtD&rdquo;: proximal-to-distal direction; For example, &ldquo;RH_Per_L_DtP.h5&rdquo; denotes a scan on the right forearm, with ultrasound probe perpendicular of the forearm sweeping along straight line, in distal-to-proximal direction.</p> </li> </ul> </li> <li> <p>Calibration matrix: The calibration matrix was obtained using a pinhead-based method. The "scaling_from_pixel_to_mm" and "spatial_calibration_from_image_coordinate_system_to_tracking_tool_coordinate_system" are provided in the &ldquo;calib_matrix.csv&rdquo;.&nbsp;</p> </li> </ul> <div> <p><strong>Data Usage Policy:</strong></p> <ul> <li>The training and validation data provided may be utilized within the research scope of this challenge and in subsequent research-related publications. However, commercial use of the training and validation data is prohibited. In cases where the intended use is ambiguous, participants accessing the data are requested to abstain from further distribution or use outside the scope of this challenge.</li> <li>If you use our dataset in your publication, please cite the challenge paper and some of the following optional articles:&nbsp; <ul> <li>Challenge paper: <ul> <li><strong>Qi Li et al. "TUS-REC2024: A Challenge to Reconstruct 3D Freehand Ultrasound Without External Tracker." <em>arXiv preprint arXiv:<a title="https://arxiv.org/abs/2506.21765" href="https://doi.org/10.48550/arXiv.2506.21765" target="_blank" rel="noopener">2506.21765</a></em>&nbsp;(2025).</strong></li> </ul> </li> <li>Optional articles: <ul> <li>Qi Li, Ziyi Shen, Qianye Yang, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Nonrigid Reconstruction of Freehand Ultrasound without a Tracker." In&nbsp;<em>International Conference on Medical Image Computing and Computer-Assisted Intervention</em>, pp. 689-699. Cham: Springer Nature Switzerland, 2024. doi: <a href="https://doi.org/10.1007/978-3-031-72083-3_64" target="_blank" rel="noopener">10.1007/978-3-031-72083-3_64.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Long-term Dependency for 3D Reconstruction of Freehand Ultrasound Without External Tracker." IEEE Transactions on Biomedical Engineering, vol. 71, no. 3, pp. 1033-1042, 2024. doi:&nbsp;<a href="https://ieeexplore.ieee.org/abstract/document/10288201" target="_blank" rel="noopener">10.1109/TBME.2023.3325551</a>.</li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Trackerless freehand ultrasound with sequence modelling and auxiliary transformation over past and future frames." In 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), pp. 1-5. IEEE, 2023. doi: <a href="https://doi.org/10.1109/ISBI53787.2023.10230773" target="_blank" rel="noopener">10.1109/ISBI53787.2023.10230773.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Privileged Anatomical and Protocol Discrimination in Trackerless 3D Ultrasound Reconstruction." In International Workshop on Advances in Simplifying Medical Ultrasound, pp. 142-151. Cham: Springer Nature Switzerland, 2023. doi: <a href="https://doi.org/10.1007/978-3-031-44521-7_14" target="_blank" rel="noopener">https://doi.org/10.1007/978-3-031-44521-7_14.</a></li> </ul> </li> </ul> </li> </ul> </div> </div> </div> </div>

opencc-by-nc-sa-4.0May 2024View details →
zenodo44/100

Trackerless 3D Freehand Ultrasound Reconstruction Challenge 2024 - Train Dataset (Part 1)

<blockquote> <p><strong>This Challenge will be an open-ended challenge, and we welcome your submission. Please register your team via this ⁠<a title="https://forms.office.com/e/dPg47ktV7M" href="https://forms.office.com/e/dPg47ktV7M" target="_blank" rel="noopener">form</a>. You can submit the algorithm via this <a title="https://forms.office.com/e/QChhNkLYiu" href="https://forms.office.com/e/QChhNkLYiu" target="_blank" rel="noopener noreferrer">form</a> for TUS-REC2024 Challenge, and we will test your submitted docker on the test set.</strong></p> <p><strong>We are organising TUS-REC2025 at MICCAI2025. More information is available on the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/" target="_blank" rel="noopener">TUS-REC2025 challenge website</a> and <a href="https://github.com/QiLi111/TUS-REC2025-Challenge_baseline" target="_blank" rel="noopener">Baseline code repo</a>.</strong></p> </blockquote> <p><strong>This is the first part of the Challenge train dataset.&nbsp;<a href="../doi/10.5281/zenodo.11180795" target="_blank" rel="noopener">Link</a> to second part; <a href="../doi/10.5281/zenodo.11355499" target="_blank" rel="noopener">Link</a> to third part. <a href="../doi/10.5281/zenodo.12979481" target="_blank" rel="noopener">Link</a> to validation dataset.</strong></p> <p>Acquisition devices and config: The 2D US images were acquired using an Ultrasonix machine (BK, Europe) with a curvilinear probe (4DC7-3/40). The associated position information of each frame was recorded by an optical tracker (NDI Polaris Vicra, Northern Digital Inc., Canada). The acquired US frames were recorded at 20 fps, with an image size of 480&times;640, without speckle reduction. The frequency was set at 6MHz with a dynamic range of 83 dB, an overall gain of 48% and a depth of 9 cm.&nbsp;</p> <div> <p>Scanning protocol: Both left and right forearms of volunteers were scanned. For each forearm, the US probe moves in three different trajectories (straight line shape, "C" shape, and "S" shape), in a distal-to-proximal direction followed by a proximal-to-distal direction, with the US plane perpendicular of and parallel to the scanning direction. The train dataset contains 1200 scans in total, 24 scans associated with each subject.</p> <p>For detailed information please refer to the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/TUS-REC2024/" target="_blank" rel="noopener">Challenge website</a>. Baseline code is also provided, which can be found at this <a href="https://github.com/QiLi111/tus-rec-challenge_baseline" target="_blank" rel="noopener">repo</a>.</p> <p>Dataset structure:&nbsp;</p> </div> <div> <ul> <li> <p>The dataset contains 50 folders (one subject per folder), each with 24 scans. Each .h5 file corresponds to one scan, storing image and transformation of each frame within this scan. Key-value pairs in each .h5 file are explained below.</p> <ul> <li> <p>&ldquo;frames&rdquo;&nbsp; - All frames in the scan; with a shape of [N,H,W], where N refers to the number of frames in the scan, H and W denote the height and width of a frame.&nbsp;</p> </li> <li> <p>&ldquo;tforms&rdquo; - All transformations in the scan; with a shape of [N,4,4], where N is the number of frames in the scan, and the transformation matrix denotes the transformation from tracker tool space to camera space.&nbsp;</p> </li> <li> <p>Notations in the name of each .h5 file: &ldquo;RH&rdquo;: right arm; &ldquo;LH&rdquo;: left arm; &ldquo;Per&rdquo;: perpendicular; &ldquo;Par&rdquo;: parallel; &ldquo;L&rdquo;: straight line shape; &ldquo;C&rdquo;: C shape; &ldquo;S&rdquo;: S shape; &ldquo;DtP&rdquo;: distal-to-proximal direction; &ldquo;PtD&rdquo;: proximal-to-distal direction; For example, &ldquo;RH_Per_L_DtP.h5&rdquo; denotes a scan on the right forearm, with ultrasound probe perpendicular of the forearm sweeping along straight line, in distal-to-proximal direction.</p> </li> </ul> </li> <li> <p>Calibration matrix: The calibration matrix was obtained using a pinhead-based method. The "scaling_from_pixel_to_mm" and "spatial_calibration_from_image_coordinate_system_to_tracking_tool_coordinate_system" are provided in the &ldquo;calib_matrix.csv&rdquo;.&nbsp;</p> </li> </ul> <div> <p><strong>Data Usage Policy:</strong></p> <ul> <li>The training and validation data provided may be utilized within the research scope of this challenge and in subsequent research-related publications. However, commercial use of the training and validation data is prohibited. In cases where the intended use is ambiguous, participants accessing the data are requested to abstain from further distribution or use outside the scope of this challenge.</li> <li>If you use our dataset in your publication, please cite the challenge paper and some of the following optional articles:&nbsp;&nbsp; <ul> <li>Challenge paper: <ul> <li><strong>Qi Li et al. "TUS-REC2024: A Challenge to Reconstruct 3D Freehand Ultrasound Without External Tracker." <em>arXiv preprint arXiv:<a title="https://arxiv.org/abs/2506.21765" href="https://doi.org/10.48550/arXiv.2506.21765" target="_blank" rel="noopener">2506.21765</a></em>&nbsp;(2025).</strong></li> </ul> </li> <li>Optional articles: <ul> <li>Qi Li, Ziyi Shen, Qianye Yang, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Nonrigid Reconstruction of Freehand Ultrasound without a Tracker." In&nbsp;<em>International Conference on Medical Image Computing and Computer-Assisted Intervention</em>, pp. 689-699. Cham: Springer Nature Switzerland, 2024. doi: <a href="https://doi.org/10.1007/978-3-031-72083-3_64" target="_blank" rel="noopener">10.1007/978-3-031-72083-3_64.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Long-term Dependency for 3D Reconstruction of Freehand Ultrasound Without External Tracker." IEEE Transactions on Biomedical Engineering, vol. 71, no. 3, pp. 1033-1042, 2024. doi:&nbsp;<a href="https://ieeexplore.ieee.org/abstract/document/10288201" target="_blank" rel="noopener">10.1109/TBME.2023.3325551</a>.</li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Trackerless freehand ultrasound with sequence modelling and auxiliary transformation over past and future frames." In 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), pp. 1-5. IEEE, 2023. doi: <a href="https://doi.org/10.1109/ISBI53787.2023.10230773" target="_blank" rel="noopener">10.1109/ISBI53787.2023.10230773.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Privileged Anatomical and Protocol Discrimination in Trackerless 3D Ultrasound Reconstruction." In International Workshop on Advances in Simplifying Medical Ultrasound, pp. 142-151. Cham: Springer Nature Switzerland, 2023. doi: <a href="https://doi.org/10.1007/978-3-031-44521-7_14" target="_blank" rel="noopener">https://doi.org/10.1007/978-3-031-44521-7_14.</a></li> </ul> </li> </ul> </li> </ul> </div> </div>

opencc-by-nc-sa-4.0May 2024View details →
zenodo44/100

Supplementary material for "Demographic assessment of reintroduced bearded vultures in the Alps: success in the core, challenges in the periphery"

<p><strong>Abstract</strong></p> <ol> <li>&nbsp;Regular assessment of reintroduced populations is essential to guide management and provide lessons for other reintroduction projects. Bearded vulture <em>Gypaetus barbatus</em> reintroduction in the Alps began in 1986 with the release of the first fledglings, the first successful reproduction was recorded in 1997, and the population has grown steadily since. A previous assessment suggested that no further releases would be required to establish a self-sustaining population from a demographic point of view. However, this conclusion was based on a small sample size and spatially homogeneous demographic rates of released individuals, which may differ from and be spatially variable among wild-hatched individuals.</li> <li>Using longitudinal data and breeding site survey data, we constructed an integrated population model to examine the demography of the entire Alpine population, spatially stratified into a core and a periphery. We performed retrospective population analyses to identify demographic reasons of spatial differences in population growth, and conducted population viability analyses to assess the impact of future threats and reintroduction options.</li> <li>In 2021, an estimated 173 (CRI: 149-199) females were present in the Alps, of which 65 (CRI: 63-67) were breeders. Adult survival and productivity were higher in the core than in the periphery, so the population grew more strongly in the core than in the periphery. Differences in adult survival contributed most to the differences in population growth between the two areas.</li> <li>The population viability analysis predicts that the Alpine population will double in 10 years but that an increase in the mortality hazard above 0.055 will lead to a population decline. Unlike the population in the core, the population in the periphery is dependent on further releases at this stage.</li> <li>Bearded vulture reintroductions in the Alps have succeeded in creating a self-sustaining population with higher reproductive success and similar survival probabilities to the autochthonous Pyrenean population. In general, management should focus on preventing further mortality risks. In the periphery, reducing current mortality and increasing reproductive success are essential to make the population independent of releases.</li> </ol> <p>&nbsp;</p> <p><strong>Readme</strong></p> <p>Data files and code for all analyses and figures presented in the paper. The two data files are provided in csv format (SightingData.csv, OccupancyData.csv). There are five code files written for R, but some the main analysis requires software JAGS. The main code (IPM_Code.txt) contains a description of the data, code for loading and managing the data, and for fitting the integrated population model. The other files contain custom written functions (Functions.txt), code for the density dependence test (DensityDependence_Code.txt), code for the retrospective analyses (tLTRE_Code.txt) and code for miscellaneous statistics and figure generation (Figure_Code.txt).</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Effects of sleep restriction and light intensity on mental effort during cognitive challenge Study 2

<p>This repository contains the data of Study 2 of the manuscript "Effects of sleep restriction and light intenisty on mental effort during cognitive challenge" by Larissa W&uuml;st and Ruta Lasauskaite.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

KGCW 2024 Challenge @ ESWC 2024

<h1><strong>Knowledge Graph Construction Workshop 2024: challenge</strong></h1> <p>Knowledge graph construction of heterogeneous data has seen a lot of uptake<br>in the last decade from compliance to performance optimizations with respect<br>to execution time. Besides execution time as a metric for comparing knowledge<br>graph construction, other metrics e.g. CPU or memory usage are not considered.<br>This challenge aims at benchmarking systems to find which RDF graph<br>construction system optimizes for metrics e.g. execution time, CPU,<br>memory usage, or a combination of these metrics.</p> <p><strong>Task description</strong></p> <p>The task is to reduce and report the execution time and computing resources<br>(CPU and memory usage) for the parameters listed in this challenge, compared<br>to the state-of-the-art of the existing tools and the baseline results provided<br>by this challenge. This challenge is not limited to execution times to create<br>the fastest pipeline, but also computing resources to achieve the most efficient<br>pipeline.</p> <p>We provide a tool which can execute such pipelines end-to-end. This tool also<br>collects and aggregates the metrics such as execution time, CPU and memory<br>usage, necessary for this challenge as CSV files. Moreover, the information<br>about the hardware used during the execution of the pipeline is available as<br>well to allow fairly comparing different pipelines. Your pipeline should consist<br>of Docker images which can be executed on Linux to run the tool. The tool is<br>already tested with existing systems, relational databases e.g. MySQL and<br>PostgreSQL, and triplestores e.g. Apache Jena Fuseki and OpenLink Virtuoso<br>which can be combined in any configuration. It is strongly encouraged to use<br>this tool for participating in this challenge. If you prefer to use a different<br>tool or our tool imposes technical requirements you cannot solve, please contact<br>us directly.</p> <h2><strong>Track 1: Conformance<br></strong></h2> <p>The set of new specification for the RDF Mapping Language (RML) established by the W3C Community Group on Knowledge Graph Construction provide a set of test-cases for each module:</p> <ul> <li> <p><a href="https://github.com/kg-construct/rml-core/">RML-Core</a></p> </li> <li> <p><a href="https://github.com/kg-construct/rml-io/">RML-IO</a></p> </li> <li> <p><a href="https://github.com/kg-construct/rml-cc/">RML-CC</a></p> </li> <li> <p><a href="https://github.com/kg-construct/rml-fnml/">RML-FNML</a></p> </li> <li> <p><a href="https://github.com/kg-construct/rml-star/">RML-Star</a></p> </li> </ul> <p>These test-cases are evaluated in this Track of the Challenge to determine their <strong>feasibility</strong>, correctness, etc. by applying them in implementations. <strong>This Track is in Beta status because these new specifications have not seen any implementation yet, thus it may contain bugs and issues.</strong> <strong>If you find problems with the mappings, output, etc. please report them to the corresponding repository of each module.</strong></p> <p><strong>Note:</strong> validating the output of the RML Star module automatically through the provided tooling is currently not possible, see <a href="https://github.com/kg-construct/challenge-tool/issues/1" target="_blank" rel="noopener">https://github.com/kg-construct/challenge-tool/issues/1</a>.</p> <p>Through this Track we aim to spark development of&nbsp; implementations for the new specifications and improve the test-cases. Let us know your problems with the test-cases and we will try to find a solution.</p> <h2><strong>Track 2: Performance</strong></h2> <p><strong>Part 1: Knowledge Graph Construction Parameters</strong></p> <p>These parameters are evaluated using synthetic generated data to have more<br>insights of their influence on the pipeline.</p> <p><em><strong>Data</strong></em></p> <ul> <li>Number of data records: scaling the data size vertically by the number of records with a fixed number of data properties (10K, 100K, 1M, 10M records).</li> <li>Number of data properties: scaling the data size horizontally by the number of data properties with a fixed number of data records (1, 10, 20, 30 columns).</li> <li>Number of duplicate values: scaling the number of duplicate values in the dataset (0%, 25%, 50%, 75%, 100%).</li> <li>Number of empty values: scaling the number of empty values in the dataset (0%, 25%, 50%, 75%, 100%).</li> <li>Number of input files: scaling the number of datasets (1, 5, 10, 15).</li> </ul> <p><em><strong>Mappings</strong></em></p> <ul> <li>Number of subjects: scaling the number of subjects with a fixed number of predicates and objects (1, 10, 20, 30 TMs).</li> <li>Number of predicates and objects: scaling the number of predicates and objects with a fixed number of subjects (1, 10, 20, 30 POMs).</li> <li>Number of and type of joins: scaling the number of joins and type of joins (1-1, N-1, 1-N, N-M)</li> </ul> <p><strong>Part 2: GTFS-Madrid-Bench</strong></p> <p>The GTFS-Madrid-Bench provides insights in the pipeline with real data from the<br>public transport domain in Madrid.</p> <p><em><strong>Scaling</strong></em></p> <ul> <li>GTFS-1 SQL</li> <li>GTFS-10 SQL</li> <li>GTFS-100 SQL</li> <li>GTFS-1000 SQL</li> </ul> <p><em><strong>Heterogeneity</strong></em></p> <ul> <li>GTFS-100 XML + JSON</li> <li>GTFS-100 CSV + XML</li> <li>GTFS-100 CSV + JSON</li> <li>GTFS-100 SQL + XML + JSON + CSV</li> </ul> <p><strong>Example pipeline</strong></p> <p>The ground truth dataset and baseline results are generated in different steps<br>for each parameter:</p> <ol> <li>The provided CSV files and SQL schema are loaded into a MySQL relational database.</li> <li>Mappings are executed by accessing the MySQL relational database to construct a knowledge graph in N-Triples as RDF format</li> </ol> <p>The pipeline is executed 5 times from which the median execution time of each<br>step is calculated and reported. Each step with the median execution time is<br>then reported in the baseline results with all its measured metrics.<br>Knowledge graph construction timeout is set to 24 hours.&nbsp;<br>The execution is performed with the following tool: <a href="https://github.com/kg-construct/challenge-tool">https://github.com/kg-construct/challenge-tool</a>,<br>you can adapt the execution plans for this example pipeline to your own needs.</p> <p>Each parameter has its own directory in the ground truth dataset with the<br>following files:</p> <ul> <li>Input dataset as CSV.</li> <li>Mapping file as RML.</li> <li>Execution plan for the pipeline in <code>metadata.json</code>.</li> </ul> <p><strong>Datasets</strong></p> <p><em><strong>Knowledge Graph Construction Parameters</strong></em></p> <p>The dataset consists of:</p> <ul> <li>Input dataset as CSV for each parameter.</li> <li>Mapping file as RML for each parameter.</li> <li>Baseline results for each parameter with the example pipeline.</li> <li>Ground truth dataset for each parameter generated with the example pipeline.</li> </ul> <p><em>Format</em></p> <p>All input datasets are provided as CSV, depending on the parameter that is being<br>evaluated, the number of rows and columns may differ. The first row is always<br>the header of the CSV.</p> <p><em><strong>GTFS-Madrid-Bench</strong></em></p> <p>The dataset consists of:</p> <ul> <li>Input dataset as CSV with SQL schema for the scaling and a combination of XML,</li> <li>CSV, and JSON is provided for the heterogeneity.</li> <li>Mapping file as RML for both scaling and heterogeneity.</li> <li>SPARQL queries to retrieve the results.</li> <li>Baseline results with the example pipeline.</li> <li>Ground truth dataset generated with the example pipeline.</li> </ul> <p><em>Format</em></p> <p>CSV datasets always have a header as their first row.<br>JSON and XML datasets have their own schema.</p> <p><strong>Evaluation criteria</strong></p> <p>Submissions must evaluate the following metrics:</p> <ul> <li>Execution time of all the steps in the pipeline. The execution time of a step is the difference between the begin and end time of a step.</li> <li>CPU time as the time spent in the CPU for all steps of the pipeline. The CPU time of a step is the difference between the begin and end CPU time of a step.</li> <li>Minimal and maximal memory consumption for each step of the pipeline. The minimal and maximal memory consumption of a step is the minimum and maximum calculated of the memory consumption during the execution of a step.</li> </ul> <p><strong>Expected output</strong></p> <p><em><strong>Duplicate values</strong></em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> </tbody> <tbody> <tr> <td>0 percent</td> <td>2000000 triples</td> </tr> <tr> <td>25 percent</td> <td>1500020 triples&nbsp;</td> </tr> <tr> <td>50 percent</td> <td>1000020 triples&nbsp;</td> </tr> <tr> <td>75 percent</td> <td>500020 triples</td> </tr> <tr> <td>100 percent</td> <td>20 triples</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><em><strong>Empty values</strong></em></p> <p>&nbsp;</p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> </tbody> <tbody> <tr> <td>0 percent</td> <td>2000000 triples</td> </tr> <tr> <td>25 percent</td> <td>1500000 triples&nbsp;</td> </tr> <tr> <td>50 percent</td> <td>1000000 triples&nbsp;</td> </tr> <tr> <td>75 percent</td> <td>500000 triples</td> </tr> <tr> <td>100 percent</td> <td>0 triples</td> </tr> </tbody> </table> <p><em><strong>Mappings</strong></em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> </tbody> <tbody> <tr> <td>1TM + 15POM</td> <td>1500000 triples</td> </tr> <tr> <td>3TM + 5POM</td> <td>1500000 triples&nbsp;</td> </tr> <tr> <td>5TM + 3POM&nbsp;</td> <td>1500000 triples&nbsp;</td> </tr> <tr> <td>15TM + 1POM</td> <td>1500000 triples</td> </tr> </tbody> </table> <p><em><strong>Properties</strong></em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> <tr> <td>1M rows 1 column</td> <td>1000000 triples</td> </tr> <tr> <td>1M rows 10 columns</td> <td>10000000 triples&nbsp;</td> </tr> <tr> <td>1M rows 20 columns</td> <td>20000000 triples&nbsp;</td> </tr> <tr> <td>1M rows 30 columns</td> <td>30000000 triples</td> </tr> </tbody> </table> <p><em><strong>Records</strong></em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> <tr> <td>10K rows 20 columns</td> <td>200000 triples</td> </tr> <tr> <td>100K rows 20 columns</td> <td>2000000 triples&nbsp;</td> </tr> <tr> <td>1M rows 20 columns</td> <td>20000000 triples&nbsp;</td> </tr> <tr> <td>10M rows 20 columns</td> <td>200000000 triples</td> </tr> </tbody> </table> <p><em><strong>Joins</strong></em></p> <p><em>1-1 joins</em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> <tr> <td>0 percent</td> <td>0 triples</td> </tr> <tr> <td>25 percent</td> <td>125000 triples&nbsp;</td> </tr> <tr> <td>50 percent</td> <td>250000 triples&nbsp;</td> </tr> <tr> <td>75 percent</td> <td>375000 triples</td> </tr> <tr> <td>100 percent</td> <td>500000 triples</td> </tr> </tbody> </table> <p><em>1-N joins</em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> <tr> <td>1-10 0 percent</td> <td>0 triples</td> </tr> <tr> <td>1-10 25 percent</td> <td>125000 triples&nbsp;</td> </tr> <tr> <td>1-10 50 percent</td> <td>250000 triples&nbsp;</td> </tr> <tr> <td>1-10 75 percent</td> <td>375000 triples</td> </tr> <tr> <td>1-10 100 percent</td> <td>500000 triples</td> </tr> <tr> <td>1-5 50 percent</td> <td>250000&nbsp;triples</td> </tr> <tr> <td>1-10 50 percent</td> <td>250000 triples&nbsp;</td> </tr> <tr> <td>1-15 50 percent</td> <td>250005 triples&nbsp;</td> </tr> <tr> <td>1-20 50 percent</td> <td>250000 triples</td> </tr> </tbody> </table> <p><em>1-N joins</em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> <tr> <td>10-1 0 percent</td> <td>0 triples</td> </tr> <tr> <td>10-1 25 percent</td> <td>125000 triples&nbsp;</td> </tr> <tr> <td>10-1 50 percent</td> <td>250000 triples&nbsp;</td> </tr> <tr> <td>10-1 75 percent</td> <td>375000 triples</td> </tr> <tr> <td>10-1 100 percent</td> <td>500000 triples</td> </tr> <tr> <td>5-1 50 percent</td> <td>250000&nbsp;triples</td> </tr> <tr> <td>10-1 50 percent</td> <td>250000 triples&nbsp;</td> </tr> <tr> <td>15-1 50 percent</td> <td>250005 triples&nbsp;</td> </tr> <tr> <td>20-1 50 percent</td> <td>250000 triples</td> </tr> </tbody> </table> <p><em>N-M joins</em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> <tr> <td>5-5 50 percent</td> <td>1374085 triples</td> </tr> <tr> <td>10-5 50 percent</td> <td>1375185 triples</td> </tr> <tr> <td>5-10 50 percent&nbsp;</td> <td>1375290 triples</td> </tr> <tr> <td>5-5 25 percent</td> <td>718785 triples</td> </tr> <tr> <td>5-5 50 percent</td> <td>1374085 triples</td> </tr> <tr> <td>5-5 75 percent&nbsp;</td> <td>1968100 triples</td> </tr> <tr> <td>5-5 100 percent&nbsp;</td> <td>2500000 triples&nbsp;</td> </tr> <tr> <td>5-10 25 percent&nbsp;</td> <td>719310 triples</td> </tr> <tr> <td>5-10 50 percent&nbsp;</td> <td>1375290 triples</td> </tr> <tr> <td>5-10 75 percent&nbsp;</td> <td>1967660 triples</td> </tr> <tr> <td>5-10 100 percent&nbsp;</td> <td>2500000 triples</td> </tr> <tr> <td>10-5 25 percent&nbsp;</td> <td>719370 triples&nbsp;</td> </tr> <tr> <td>10-5 50 percent&nbsp;</td> <td>1375185 triples</td> </tr> <tr> <td>10-5 75 percent&nbsp;</td> <td>1968235 triples</td> </tr> <tr> <td>10-5 100 percent&nbsp;</td> <td>2500000 triples</td> </tr> </tbody> </table> <p><strong>GTFS Madrid Bench</strong></p> <p><em>Generated Knowledge Graph</em></p> <table> <tbody> <tr> <th>Scale</th> <th>Number of Triples</th> </tr> <tr> <td>1</td> <td>395953 triples</td> </tr> <tr> <td>10</td> <td>3959530 triples&nbsp;</td> </tr> <tr> <td>100</td> <td>39595300 triples&nbsp;</td> </tr> <tr> <td>1000</td> <td>395953000 triples</td> </tr> </tbody> </table> <p><em>Queries</em></p> <table> <tbody> <tr> <th>Query</th> <th>Scale 1</th> <th>Scale 10</th> <th>Scale 100</th> <th>Scale 1000</th> </tr> <tr> <td>Q1</td> <td>58540 results</td> <td>585400 results</td> <td>No results available</td> <td>No results available</td> </tr> <tr> <td>Q2</td> <td>636 results</td> <td>11998 results&nbsp;&nbsp;</td> <td>125565 results</td> <td>1261368 results</td> </tr> <tr> <td>Q3</td> <td>421 results</td> <td>4207 results&nbsp;</td> <td>42067 results</td> <td>420667 results</td> </tr> <tr> <td>Q4</td> <td>13 results</td> <td>130 results</td> <td>1300 results</td> <td>13000 results</td> </tr> <tr> <td>Q5</td> <td>35 results</td> <td>350 results</td> <td>3500 results</td> <td>35000 results</td> </tr> <tr> <td>Q6</td> <td>1 result</td> <td>1 result</td> <td>1 result</td> <td>1 result</td> </tr> <tr> <td>Q7</td> <td>68 results</td> <td>67 results</td> <td>67 results</td> <td>53 results</td> </tr> <tr> <td>Q8</td> <td>35460 results</td> <td>354600 results</td> <td>No results available</td> <td>No results available</td> </tr> <tr> <td>Q9</td> <td>130 results</td> <td>1300 results</td> <td>13000 results</td> <td>130000 results</td> </tr> <tr> <td>Q10</td> <td>1 result</td> <td>1 result</td> <td>1 result</td> <td>1 result</td> </tr> <tr> <td>Q11</td> <td>130 results</td> <td>260 results&nbsp;</td> <td>260 results&nbsp;</td> <td>260 results&nbsp;</td> </tr> <tr> <td>Q12</td> <td>13 results</td> <td>130 results</td> <td>1300 results</td> <td>13000 results</td> </tr> <tr> <td>Q13</td> <td>265 results</td> <td>2650 results</td> <td>26500 results</td> <td>265000 results</td> </tr> <tr> <td>Q14</td> <td>2234 results</td> <td>22340 results</td> <td>223400 results</td> <td>No results available</td> </tr> <tr> <td>Q15</td> <td>592 results</td> <td>8684 results</td> <td>35502 results</td> <td>206628 results</td> </tr> <tr> <td>Q16</td> <td>390 results</td> <td>780 results</td> <td>260 results</td> <td>780 results</td> </tr> <tr> <td>Q17</td> <td>855 results</td> <td>8550 results&nbsp;</td> <td>85500 results&nbsp;</td> <td>855000 results&nbsp;</td> </tr> <tr> <td>Q18</td> <td>104 results</td> <td>1300 results</td> <td>13000 results</td> <td>130000 results</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

A Patent Survey on the Biotechnological Production of 2,5-Furandicarboxylic acid (FDCA): Current Trends and Challenges

<p>The production of 2,5-furandicarboxylic acid (FDCA) as a biobased commodity chemical has gained great importance over the last years. The possibility of replacing conventional polyethylene terephthalate (PET)-based plastics by polyethylene furanoate (PEF) is accelerating applied research in the direction of novel and highly productive synthesis routes to FDCA. This paper explores the patent activities related to FDCA production, with particular emphasis on the potential role that enzymatic catalytic methods may have. An increasing number of patent applications (and granted ones) have been disclosed over the last decade. The innovation is set on the development of multi-step catalytic processes, involving multi-enzymatic and chemoenzymatic cascades, or using whole-cells, either as resting (non-growing) biocatalysts or as living organisms in fermentative processes. Moreover, other innovative paths use substrates different from fructose and 5-hydroxymethylfurfural (HMF), such as gluconic acid, or furfural. Overall, opportunities for innovation exist, albeit current production metrics in biotechnological methods remain mostly at the proof-of-concept level. Large development efforts to reach industrial targets are needed, which may be stimulated by the less severe processing conditions expected for biotechnology, leading to energy savings, less by-product formation, and to the valorization of crude effluents.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

The evolution of gender monitoring and its challenges in Research and Innovation in Europe: She Figures reports analysis dataset

<p><span>The article delves into the European Commission's flagship initiative on gender monitoring in science and innovation, offering a responsible metrics perspective and drawing on equality policy literature. Over two decades, the initiative has evolved from competitiveness-related justifications to more transformative objectives related to equality policy evaluation, with the measurement areas and policy focus also undergoing changes. While there has been notable progress, the article points out a logic of invisibility in how dimensions and indicators are conceptualised and their data sources and interpretation. However, it also highlights a significant improvement in the information available. The article suggests that the contextualisation of the process could be enhanced to better integrate it into the policy-making cycle, a crucial area for further research. It concludes with proposals for future gender monitoring science and innovation. The aim is to offer an encouraging vision of monitoring that counts more on who is monitored and in opening up the debates instead of closing them.</span></p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

CryoEM Models and Associated Data Submitted to the 2015/2016 EMDataBank Model Challenge

<p>Files and metadata associated with the EMDataBank/Unified Data Resource for 3DEM 2015/2016 Models Challenge hosted at challenges.emdatabank.org are deposited.</p> <p>All members of the Scientific Community--at all levels of experience--were invited to participate as Challengers, and/or as Assessors.</p> <p>Eight recently determined target structures were selected for the challenge. All of the maps were archived in the EM Data Bank (EMDB; http://emdatabank.org).</p> <p>In total 16 Challengers created 106 models and uploaded their results with associated details.&nbsp; In the zip files, each entry is represented in a folder containing the original deposition upload (deposited_EM.pdb), initial processing at RCSB/Rutgers (deposited_EM_edited.pdb, maxit.cif, maxit.cif.pdb) and final model version evaluated (model-compare.pdb) at UC Davis (http://model-compare.emdatabank.org).</p> <p>This model challenge was one of two community-wide challenges sponsored by EMDataBank in 2015/2016 to critically evaluate 3DEM methods that are coming into use, with the ultimate goal of developing validation criteria associated with every 3DEM map and map-derived model.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2018View details →
zenodo44/100

PlantVillage Disease Classification Challenge - Color Images

<p><br> This work is licensed under a <a href="http://creativecommons.org/licenses/by-sa/3.0/us/">Creative Commons Attribution-ShareAlike 3.0 United States License</a>.<br> <br> # Data origins<br> The dataset is originally hosted at <a href="https://www.crowdai.org/challenges/plantvillage-disease-classification-challenge">PlantVillage Disease Classification Challenge</a>.<br> We use the modified version in <a href="https://github.com/salathegroup/plantvillage_deeplearning_paper_dataset">this github repository</a> to do controlled experiments.<br> We only use the raw color images dataset and delete the unconventional characters in the classes directory name and `.csv` filenames.<br> <br> # Directory explanation<br> The `80-20` direcotry has multiple `.txt` files which contain the training (~80%), validation(~10%) and testing (~10%) datasets instances filenames and the corresponding label indexes. The validation dataset quantity is `5430` in all data separation. In our experiment code (not included in this archive), the validation and testing dataset are merged together.<br> <br> # Data usage<br> ## Replicate our experiments<br> We have used this dataset in writing our paper. The reference information can be seen at https://<a href="https://gitlab.com/huix/leaf-disease-plant-village">gitlab.com/huix/leaf-disease-plant-village</a>.<br> <br> ### Steps<br> 1. `cd` to the direcotry (e.g. `/home/usrname/plantvillage_deeplearning_paper_dataset`) that contains the `color` directory.<br> 2. run `python change_filename_prefix.py --prefix /home/usrname/plantvillage_deeplearning_paper_dataset` to modify the prefix path (which is `/home/h/plantvillage_deeplearning_paper_dataset` in our former generated datasets).<br> 3. Fin. You can use our <a href="https://gitlab.com/huix/leaf-disease-plant-village">opens ource codes repository</a> to do the later experiments.<br> <br> ## Generate your own training/validation/testing datasets<br> This data separation generating code isn&#39;t included in the dataset archive, it is in our open source code. Please see our <a href="https://gitlab.com/huix/leaf-disease-plant-village">open source code repository</a> for the detailed information.<br> If you have any questions, you can contact the author through email.<br> The email address is a QR code in the archive.</p>

opencc-by-4.0Mar 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record