Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

28,650

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

28,650 results for “trial”

Learn how ShareScore rates datasets ↗
zenodo40/100

data for "Plant genetic effects on microbial hubs impact host fitness in repeated field trials"

<p>These are data tables required for the analysis of the paper &quot;Plant genetic effects on microbial hubs impact host fitness in repeated field trials&quot;.&nbsp;<br> All scripts are available at&nbsp;https://forgemia.inra.fr/bbrachi/microbiota_paper.git</p> <p>The folder architecture in the zip files is the same as in the repository:&nbsp;https://forgemia.inra.fr/bbrachi/microbiota_paper.git</p> <p>The dataset includes:&nbsp;</p> <p>- OTU count tables for 16S and ITS</p> <p>- taxonomic assignation</p> <p>- plant seed-set estimates</p> <p>- plant growth data from the B38 experiment.&nbsp;</p> <p>- Metabolomics datasets</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

A multicenter randomized phase 4 trial comparing Sodium Picosulphate plus Magnesium Citrate vs Polyethylene Glycol plus Ascorbic Acid for bowel PREparation before COLonoscopy. The PRECOL trial.

<p>This is the database for final analysis of the PRECOL clinical trial, whose abstract follows</p> <p>Background<strong>.</strong> Adequate bowel preparation before colonoscopy is crucial. Unfortunately, up to 25% of all colonoscopies have inadequate bowel cleansing.&nbsp; From a patient perspective, bowel preparation is a main obstacle for colonoscopy. Several low-volume bowel preparations have been formulated to provide more tolerable purgative solutions without loss of efficacy.</p> <p>Methods: In this phase 4, randomized, multicenter, two-arm trial, adult outpatients undergoing colonoscopy received either Sodium Picosulphate plus Magnesium Citrate (SPMC) or Polyethylene Glycol plus Ascorbic Acid (PEG-ASC) for bowel preparation. The primary aims were to test quality of bowel cleansing (primary endpoint, scored according the Boston Bowel Preparation Scale) and patient&rsquo;s acceptance (measured with 6 visual analogue scales). The study was open as for treatment assignment, and blinded for primary endpoint assessment that was done independently on videotaped colonoscopies by 2 endoscopists not aware of study arm. A sample size of 525 patients was calculated to recognize a difference of 10% in the proportion of successes between the arms with a two-sided alpha error of 0&middot;05 and 90% statistical power.</p> <p>Findings: overall 550 subjects (279 assigned to PEG-ASC and 271 assigned to SPMC) represented the analysis population. There was no statistically significant difference in the success rate according to BBPS: 94&middot;4% with PEG-ASC and 95&middot;7% with SPMC (P=0&middot;49). Acceptance and willing to repeat were significantly better for SPMC with all the scales. Compliance was less than full in 6&middot;6% and 9&middot;9% of cases with PEG-ASC and SPMC, respectively (P=0&middot;17). Nausea and meteorism were significantly more bothersome with PEG-ASC than SPMC. There were no serious adverse events in either group.</p> <p>Interpretation. SPMC and PEG-ASC are not different in terms of efficacy, but SPMC is better tolerated than PEG-ASC. SPMC could be used as alternative to low-volume PEG based purgative solutions for bowel preparation.</p> <p>Funding. This research had no financial support.</p> <p>ClinicalTrials.gov NCT01649674; EudraCT 2011&mdash;000587&mdash;10.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Classification of hierarchical text using geometric deep learning: the case of clinical trials corpus

<p>We consider the hierarchical representation of documents as graphs and use geometric deep learning to classify them into different categories. While graph neural networks can efficiently handle the variable structure of hierarchical documents using the permutation invariant message passing operations, we show that we can gain extra performance improvements using our proposed selective graph pooling operation that arises from the fact that some parts of the hierarchy are invariable across different documents. We applied our model to classify clinical trial (CT) protocols into completed and terminated categories. We use bag-of-words based as well as pre-trained transformer-based embeddings to featurize the graph nodes, achieving f1-scores $\simeq 0.85$ on a publicly available large scale CT registry of around 360K protocols. We further demonstrate how the selective pooling can add insights into the CT termination status prediction.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Pre- and post-intervention responses to a knowledge, attitudes, and practices survey for the study, "Disseminating vaccination information in baby soap products increases knowledge and vaccine uptake in central Uganda: A non-randomized controlled trial"

<p>This dataset contains responses to the&nbsp;pre- and post-intervention&nbsp;knowledge, attitudes, and practices surveys utilized&nbsp;for the study, &quot;Disseminating vaccination information in baby soap products increases knowledge and vaccine uptake in central Uganda: A non-randomized controlled trial.&quot;</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Data presented in Multi-trial analysis of HIV-1 envelope gp41-reactive antibodies among global recipients of candidate HIV-1 vaccines.

<p>This folder contains datasets analyzed&nbsp;in the manuscript:</p> <p>Multi-trial analysis of HIV-1 envelope gp41-reactive antibodies among global recipients of candidate HIV-1 vaccines.</p> <p>Frontiers&nbsp;in Immunology<br> Sec. Vaccines and Molecular Therapeutics<br> doi: 10.3389/fimmu.2022.983313</p>

opencc-by-4.0Dec 2021View details →
dryad40/100

Feasibility and acceptability of personalized breast cancer screening (DECIDO Study): A single-arm proof-of-concept trial

<p>The aim of this study was to assess the acceptability and feasibility of offering risk-based breast cancer screening and its integration into regular clinical practice. A single-arm proof-of-concept trial was conducted with a sample of 387 women aged 40–50 years residing in the city of Lleida (Spain). The study intervention consisted of breast cancer risk estimation, risk communication and screening recommendations, and a follow-up. A polygenic risk score with 83 single nucleotide polymorphisms was used to update the Breast Cancer Surveillance Consortium risk model and estimate the 5-year absolute risk of breast cancer. The women expressed a positive attitude towards varying the frequency of breast screening according to individual risk and, especially, more frequently inviting women at higher-than-average risk. A lower intensity screening for women at lower risk was not as welcome, although half of the participants would accept it. Knowledge of the benefits and harms of breast screening was low, especially with regard to false positives and overdiagnosis. The women expressed a high understanding of individual risk and screening recommendations. The participants' intention to participate in risk-based screening and satisfaction at 1-year were very high.</p>

opencc-zeroSep 2022View details →
zenodo40/100

Animal and pasture data of AGROMIX WP3 pilot site trial at Tenuta di Paganico (GR) Italy (Spring 2021)

<p>Dataset of collected data on herbage, production , animal intake, and animal welfare in spring 2021 at Tenuta di Paganico (GR) Italy</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Supplementary material 1 from: Motloung R, Robertson M, Rouget M, Wilson J (2014) Forestry trial data can be used to evaluate climate-based species distribution models in predicting tree invasions. NeoBiota 20: 31-48. https://doi.org/10.3897/neobiota.20.5778

Current and potential distributions of sixteen species that are not widespread in southern Africa arranged on the basis of their suitable range size : a) Acacia paradoxa, b) A. cultriformis, c) A. falciformis, d) A. pendula, e) A. rubida, f) A. stricta, g) A. retinodes, h) A. fimbriata, i) A. aneura, j) A. viscidula, k) A. acuminata, l) A. adunca, m) A. binervata, n) A. schinoides, o) A. prominens, p) A. mangium. The grey shading indicates areas that SDMs have identified as suitable by SDMs while the white ones are unsuitable.

opencc-by-4.0Jan 2014View details →
zenodo40/100

Dataset supplementing the article A randomised controlled trial of Losartan as an anti-fibrotic agent in non-alcoholic steatohepatitis

<p>These data supplement the article A randomised controlled trial of Losartan as an antifibrotic agent in non-alcoholic steatohepatitis.</p> <p>Stuart McPherson, Nina Wilkinson, Dina Tiniakos, Jennifer Wilkinson, Alastair Burt, Elaine McColl, Deborah D. Stocken, Nick Steen, Jane Barnes, Nicola Goudie, Stephen Stewart, Yvonne Bury5, Derek Mann, Quentin M. Anstee, Christopher P. Day.</p> <p>Use is free for academic purposes, provided the aforementioned article is appropriately cited.</p> <p>The data consists of anonymised csv data files, described in DataDictionaryFELINE.csv, the final protocol and a blank copy of the case report forms used (both pdf).</p>

opencc-by-nc-4.0Dec 2016View details →
zenodo40/100

MultiDEFusion trial repository

<h1>Instructions for downloading a trial repository for the MultiDEFusion library</h1> <p>The following repository has been created as a trial dataset for the MultiDEFusion library.</p> <p>The dataset can be downloaded from GitHub or Zenodo platform.</p> <h2>Cloning the repository from GitHub</h2> <ol> <li>Open a terminal or command prompt.</li> <li>Use the git clone command to clone the repository to your device:<br><code>git clone https://github.com/damiantondas/multidefusion_trial.git</code></li> <li>The repository will be downloaded to the current directory. You can now navigate to the repository directory using the <code>cd</code> command:&nbsp;<code>cd multidefusion_trial</code></li> </ol> <h2>Cloning the repository from Zenodo</h2> <ol> <li>Download the <strong>multidefusion_trial.zip</strong> folder.</li> <li>Unzip the folder.</li> </ol> <h2>Running the integration procedure</h2> <ol> <li>To run the integration procedure in the Python environment, the initial parameters are required to be defined by the user. In the following, you can find an example script to run fusion for <code>ALL</code> stations in <code>multidefusion_trial</code> folder using <code>forward-backward</code> method with <code>0.03</code> mm/day2 noise level:</li> </ol> <p><code>from multidefusion.fusion import run_fusion</code></p> <p><code>integration = run_fusion(stations="ALL", path="/path/to/multidefusion_trial/", method="forward-backward", noise=0.03)</code></p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 2.&nbsp; More examples can be found in the <a href="https://damiantondas.github.io/multidefusion/usage/">library documentation</a>.</p>

openmit-licenseApr 2024View details →
dryad40/100

Randomized controlled oncology trials with tumor stage inclusion criteria

<p><em>Background:</em></p> <p>Extracting inclusion and exclusion criteria in a structured, automated fashion remains a challenge to developing better search functionalities or automating systematic reviews of randomized controlled trials in oncology. The question "Did this trial enroll patients with localized disease, metastatic disease, or both?" could be used to narrow down the number of potentially relevant trials when conducting a search.</p> <p><em>Dataset collection:</em></p> <p>600 randomized controlled trials from high-impact medical journals were classified depending on whether they allowed for the inclusion of patients with localized and/or metastatic disease. The dataset was randomly split into a training/validation and a test set of 500 and 100 trials respectively. However, the sets could be merged to allow for different splits.</p> <p><em>Data properties:</em></p> <p>Each trial is a row in the csv file. For each trial there is a doi, a publication date, a title, an abstract, the abstract sections (introduction, methods, results, conclusion), several tags associated with the annotation process (text, _input_hash, _task_hash, options, _view_id, config, accept, answer, _timestamp, _annotator_id,_session_id), and the assigned labels (answer).</p>

opencc-zeroJun 2024View details →
zenodo40/100

Effects of personalized music listening on post-stroke cognitive impairment: A randomized controlled trial

<p><strong><span>Background and purpose:</span></strong><span> Previous studies have suggested that music listening has the potential to positively affect mood and cognitive functions in individuals with <a name="_Hlk140153246"></a>post-stroke cognitive impairment (PSCI), with a preference for self-selected music likely to yield better outcomes. However, there is insufficient clinical evidence to suggest the use of music listening in routine rehabilitation care to treat PSCI. This randomized control trial (RCT) aims to investigate the effects of personalized music listening on mood improvement, <a name="_Hlk140153263"></a>activities of daily living (ADLs), and cognitive functions in individuals with PSCI.</span></p> <p><strong><span>Materials and methods:</span></strong><span> A total of 34 patients with PSCI were randomly assigned to either the music group or the control group. Patients in the music group underwent a three-month personalized music-listening intervention. The intervention involved listening to a personalized playlist tailored to each individual's cultural, ethnic, and social background, life experiences, and personal music preferences. In contrast, the control group patients listened to white noise as a placebo. Cognitive function, neurological function, mood, and ADLs were assessed. </span></p> <p><strong><span>Results:</span></strong><strong><span> </span></strong><span>After three months of treatment, the music group showed significantly higher <a name="_Hlk140153278"></a>Montreal Cognitive Assessment (MoCA) scores compared to the control group (<em>p=</em>0.027), particularly in the domains of delayed memory (<em>p=</em>0.019) and orientation (<em>p=</em>0.023). Moreover, the music group demonstrated significantly better scores in <a name="_Hlk140153297"></a>National Institute of Health Stroke Scale (NIHSS) (<em>p=</em>0.008), <a name="_Hlk140153304"></a>Barthel Index (BI) (<em>p=</em>0.019), and <a name="_Hlk140153315"></a>Zarit Caregiver Burden Interview (ZBI) (<em>p=</em>0.008) compared to the control group. No effects were found on mood as measured by the Hamilton Rating Scale for Anxiety (HAMA) and the Hamilton depression scale (HAMD).</span></p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 2. Representative trial format for showing side B in Avoidance of cold-, cool-, and warm-water fishes to Zequanox exposure

Figure 2. Representative trial format for showing side B (above centerline) treated first. Zequanox concentrations in mg/L as active ingredient (shaded area) for each side of the choice tank. One trial per species (n = 6) was sampled every 5 minutes during the control period and every 2 minutes thereafter. The control period was preceded by a 10-minute acclimation period.

opencc-by-4.0Oct 2020View details →
zenodo40/100

Datasets: Laser scarecrows reduce avian corn-foraging propensity but not bout length in aviary trials

<p>This archive is comprised of 3 files:</p> <p>(1) Archive Metadata: a description of the data collection, behavioral sampling, and datafile structure (variables);</p> <p>(2) An excel file containing scan sample data used in 2 analyses; and&nbsp;</p> <p>(3) An excel file containing focal foraging bout data for a 3rd analysis for the named manuscript.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Multilingual Speaker Anonymization Trials for CommonVoice and Multilingual LibriSpeech

<p>This dataset contains the speaker verification trial files of the evaluation data splits for Multilingual LibriSpeech (MLS) and CommonVoice (CV) that we propose in our paper "Probing the Feasibility of Multilingual Speaker Anonymization". <strong>The actual audio files are not included and have to be obtained separately</strong>, following the licenses of the respective corpus creators. The files in this dataset only contain the audio file IDs that can be used to prepare the data for the evaluation.&nbsp;</p> <h2>Data</h2> <p>All files contain the utterance IDs of the original MLS and CV corpora which makes it possible to align them to the audio files as provided by the dataset creators. If you use these datasets, you need to cite the original sources. We do not claim any rights to the audios.</p> <p>The dataset does not include all IDs of the corpora. For MLS, "dev" and "test" correspond to the dev and test splits as provided in the MLS corpus. For CV,&nbsp; the corpus was divided randomly into dev and test while ensuring no speaker overlap between both splits. We use the CV 16.1 version of the corpus and restrict it to validated audios where the user specified their gender as either female or male. Further information about the data restrictions can be found in the paper.</p> <p>Please note that we use the "client ID" of CV to distinguish between speakers. We acknowledge that this can lead to having the same speaker under different name multiple times in the corpus if they were assigned different client IDs. Based on our results for the original (non-anonymized) data, we believe this to have only little effect on our evaluation data.</p> <p>The following languages are included in this dataset: <strong>English (en), German (de), Dutch (nl), French (fr), Spanish (es), Italian (it), Portuguese (pt), Polish (pl), and Russian (ru).</strong></p> <h2>Structure</h2> <p>The directory contains separate folders for MLS and CV, with subfolders for each language. Each language subfolder contains 7 files:</p> <ul> <li>dev_enrolls</li> <li>dev_trials_f</li> <li>dev_trials_m</li> <li>test_enrolls</li> <li>test_trials_f</li> <li>test_trials_m</li> <li>utt2spk</li> </ul> <p>The file structures follow the evaluation data of the Voice Privacy Challenges (<a href="https://www.voiceprivacychallenge.org" rel="nofollow">https://www.voiceprivacychallenge.org</a>). The "enrolls" file contain the list of utterance IDs (i.e., audio files) that are used for enrollment of the speaker verification model. "trials_f" and "trials_m" correspond to the trial files for female and male speakers, respectively. Each line in a trial file consists of three constituents, separated by space: "enrollment speaker" "trial utterance" "target/nontarget". The last constituent signals whether the trial utterance was originally (i.e., before the anonymization) spoken by the enrollment speaker (target) or not (nontarget). The utt2spk file contains the true mapping between utterance and original speaker. This file is especially important for the CV corpus where we created new speaker names to replace the long client IDs, and where the speaker assignment is not visible in the file name.</p> <h2>Creation Process</h2> <p>For the preparation of the data into enrollment and trial subset, we tried to follow the dev and test files of the Voice Privacy Challenge 2022. This results in far more nontarget than target trials, and additional speakers in the trial set that are not contained in the enrollment set.</p> <h3>MLS</h3> <p>The MLS corpus (<a href="https://www.openslr.org/94/" rel="nofollow">available here</a>) already comes with a split into train / dev / test, which we reuse here. The data is further divided into 8 languages: en, de, nl, fr, es, it, pt and pl. We only take the dev and test sets for 6 languages (de, nl, fr, es, it, and pt). For en, we already have an alternative from the Voice Privacy Challenges based on the monolingual English LibriSpeech, so there is no need for another part from MLS. For pl, the MLS part only contains 2 speakers per gender and dev / test split, which is too small for effective speaker verification. <em>(Sidenote: Dutch is not significantly bigger with only 3 speakers per gender and split, but we decided to keep it anyway).</em> We do not further restrict the number of utterances or speaker in each language and dev / test set, which results in a large inbalance across languages. However, as given in the original MLS splits, the languages itself are balanced in terms of gender.</p> <h3>CV</h3> <p>Mozilla's CommonVoice data collection is significantly bigger than MLS, so we could select more speakers per language and make sure that the datasets per language were more or less balanced. We take the data from CV Version 16.1 (<a href="https://commonvoice.mozilla.org/en/datasets" rel="nofollow">available here</a>). Please note that users who <em>donated</em> their voice to the data collection can opt out of being included in the data at any point. <strong>This means that some speakers or utterances contained in our trial data might be missing in future downloads of the CV corpus.</strong> We further want to mention that we had to use the field <code>client ID</code> in the CV corpus for speaker assignment which is not fully accurate. The same speaker might end up with several client IDs if they are recording the utterances in different sessions. However, for the purpose of speaker anonymization, this issue is not as relevant as for pure speaker recognition.</p> <p>We use CV for all of our 9 languages. CV does not come with a division into train / dev / test splits, so we randomly sample speakers and utterances from it. For this sampling, we consider only speakers that have gender annotated as either female or male, and have recorded at least 50 validated audios. We further make sure that we have the same number of female as male speakers which leads for several languages to a large reduction in size, with most languages having far more male speakers in CV than female speakers. This leads especially to a smaller dataset for pl, for which only 14 speakers per gender and dev / test split are available. We randomly select at most 20 speakers per gender for each dev and test in each language, and randomly choose up to 70 utterances per speaker. As our lower bound for speaker selection was originally 50 utterances per speaker, this results in 50-70 utterances per speaker.</p> <h3>Separation into Enrollment and Trials</h3> <p>In MLS, we use all speakers for the enrollment and trial. The only exception is de, for which more speakers are available. In the MLS-de data, we reserve 5 speakers per gender and split for trial only, which creates some unseen distraction speakers in the trial data. In CV, we use 15 speakers per gender and split for enrollment (except for the smaller pl part, for which it is only 10), and also reserve up to 5 speakers per gender and split only for trial. Naturally, all enrollment speakers are also used in trial.</p> <p>15% of all utterances of a speaker (at least 5 utterances) are used as enrollment utterances, the rest for trial. All trial utterances are paired with each enrollment speaker of the respective gender. If an enrollment speaker is the actual speaker of that utterance, this is denoted as <em>target</em>, otherwise as <em>nontarget</em>.</p> <p>During the trials, the enrollment speaker is modeled as an average of the speaker embeddings of all enrollment utterances of that speaker.</p> <h2>Statistics</h2> <p>The following section displays the statistics for each dataset, language and dev / test split. In these statistics, female and male speakers are not distinguished, but the numbers are balanced for each subset.</p> <p>The following information is given for each dataset and language:</p> <ul> <li># speakers: number of speakers used for both enrollment and trial (50% female / 50% male)</li> <li># add.trial speakers: number of speakers additionally used only in trial</li> <li># enroll utts: total number of utterances used in enrollment (across all speakers)</li> <li># trial utts: total number of utterances used in trials (across all speakers)</li> <li># target trials: number of target trials (enrollment speaker == trial speaker)</li> <li># nontarget trials: number of nontarget trials (enrollment speaker != trial speaker)</li> <li># words: total number of words across all trial utterances (the WER is computed based on them)</li> <li># avg. utt length: average length of all utterances in the dataset, in seconds</li> </ul> <div> <h3>Development Data</h3> <p><strong>Total dataset statistics:</strong></p> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># speakers</th> <th># add. trial speakers</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> <th>avg. utt length</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>20</td> <td>10</td> <td>333</td> <td>3,136</td> <td>1,936</td> <td>29,424</td> <td>111,245</td> <td>14.90</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>18</td> <td>0</td> <td>354</td> <td>2,062</td> <td>2,062</td> <td>16,496</td> <td>73,007</td> <td>15.03</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>10</td> <td>0</td> <td>183</td> <td>1,065</td> <td>1,065</td> <td>4,260</td> <td>34,636</td> <td>14.88</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>20</td> <td>0</td> <td>349</td> <td>2,059</td> <td>2,059</td> <td>18,531</td> <td>74,782</td> <td>14.95</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>10</td> <td>0</td> <td>119</td> <td>707</td> <td>707</td> <td>2,828</td> <td>24,733</td> <td>15.90</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>6</td> <td>0</td> <td>461</td> <td>2,634</td> <td>2,634</td> <td>5,268</td> <td>11,0384</td> <td>14.83</td> </tr> <tr> <td>CV</td> <td>en</td> <td>30</td> <td>10</td> <td>279</td> <td>2,306</td> <td>1,691</td> <td>32,899</td> <td>22,394</td> <td>4.96</td> </tr> <tr> <td>&nbsp;</td> <td>de</td> <td>30</td> <td>9</td> <td>289</td> <td>2,396</td> <td>1,738</td> <td>34,202</td> <td>21,421</td> <td>5.16</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>30</td> <td>10</td> <td>293</td> <td>2,432</td> <td>1,761</td> <td>34,719</td> <td>23,240</td> <td>4.93</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>30</td> <td>10</td> <td>286</td> <td>2,421</td> <td>1,722</td> <td>34,593</td> <td>23,745</td> <td>5.69</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>30</td> <td>10</td> <td>290</td> <td>2,401</td> <td>1,735</td> <td>34,280</td> <td>22,569</td> <td>5.23</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>30</td> <td>10</td> <td>284</td> <td>2,398</td> <td>1,708</td> <td>34,262</td> <td>16,411</td> <td>3.97</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>30</td> <td>10</td> <td>288</td> <td>2,401</td> <td>1,725</td> <td>34,290</td> <td>22,108</td> <td>4.51</td> </tr> <tr> <td>&nbsp;</td> <td>pl</td> <td>20</td> <td>8</td> <td>199</td> <td>1,739</td> <td>1,196</td> <td>16,194</td> <td>12,876</td> <td>4.39</td> </tr> <tr> <td>&nbsp;</td> <td>ru</td> <td>30</td> <td>10</td> <td>289</td> <td>2,429</td> <td>1,737</td> <td>34,698</td> <td>20,599</td> <td>5.24</td> </tr> </tbody> </table> <div> <p><strong>Dataset statistics per speaker (average)</strong></p> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>16.6</td> <td>104.5</td> <td>96.8</td> <td>1,471.2</td> <td>3,553.0</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>19.7</td> <td>114.6</td> <td>114.6</td> <td>916.4</td> <td>4,350.0</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>18.3</td> <td>106.5</td> <td>106.5</td> <td>426.0</td> <td>4,405.5</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>17.4</td> <td>103.0</td> <td>103.0</td> <td>926.6</td> <td>4,300.5</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>11.9</td> <td>70.7</td> <td>70.7</td> <td>282.8</td> <td>2,084.0</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>76.8</td> <td>439.0</td> <td>439.0</td> <td>878.0</td> <td>45,515.0</td> </tr> <tr> <td>CV</td> <td>en</td> <td>9.3</td> <td>59.1</td> <td>56.4</td> <td>1,096.6</td> <td>627.0</td> </tr> <tr> <td>&nbsp;</td> <td>de</td> <td>9.6</td> <td>59.9</td> <td>57.9</td> <td>1,140.1</td> <td>492.5</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>9.8</td> <td>60.8</td> <td>58.7</td> <td>1,157.3</td> <td>606.0</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>9.5</td> <td>60.5</td> <td>57.4</td> <td>1,153.1</td> <td>594.5</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>9.7</td> <td>60.0</td> <td>57.8</td> <td>1,142.7</td> <td>494.0</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>9.5</td> <td>60.0</td> <td>56.9</td> <td>1,142.1</td> <td>308.0</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>9.6</td> <td>60.0</td> <td>57.5</td> <td>1,143.0</td> <td>520.5</td> </tr> <tr> <td>&nbsp;</td> <td>pl</td> <td>10.0</td> <td>62.1</td> <td>59.8</td> <td>809.7</td> <td>514.5</td> </tr> <tr> <td>&nbsp;</td> <td>ru</td> <td>9.6</td> <td>60.7</td> <td>57.9</td> <td>1,156.6</td> <td>502.0</td> </tr> </tbody> </table> <h3>Test Data</h3> <p><strong>Total dataset statistics:</strong></p> <p>&nbsp;</p> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># speakers</th> <th># add. trial speakers</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> <th>avg. utt length</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>30</td> <td>0</td> <td>329</td> <td>3,065</td> <td>1,906</td> <td>28,744</td> <td>110,202</td> <td>15.18</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>18</td> <td>0</td> <td>357</td> <td>2,069</td> <td>2,069</td> <td>16,552</td> <td>79,524</td> <td>14.94</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>10</td> <td>0</td> <td>185</td> <td>1,077</td> <td>1,077</td> <td>4,308</td> <td>34,796</td> <td>15.07</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>20</td> <td>0</td> <td>348</td> <td>2,037</td> <td>2,037</td> <td>18,333</td> <td>75,536</td> <td>15.11</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>10</td> <td>0</td> <td>125</td> <td>746</td> <td>746</td> <td>2,984</td> <td>26,769</td> <td>15.47</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>6</td> <td>0</td> <td>458</td> <td>2,617</td> <td>2,617</td> <td>5,234</td> <td>108,489</td> <td>14.96</td> </tr> <tr> <td>CV</td> <td>en</td> <td>30</td> <td>9</td> <td>289</td> <td>2,344</td> <td>1,733</td> <td>33,427</td> <td>22,560</td> <td>5.09</td> </tr> <tr> <td>&nbsp;</td> <td>de</td> <td>30</td> <td>10</td> <td>284</td> <td>2,377</td> <td>1,713</td> <td>33,942</td> <td>21,492</td> <td>5.24</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>30</td> <td>10</td> <td>291</td> <td>2,408</td> <td>1,745</td> <td>34,275</td> <td>22,887</td> <td>4.74</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>30</td> <td>10</td> <td>298</td> <td>2,444</td> <td>1,786</td> <td>34,874</td> <td>24,091</td> <td>5.37</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>30</td> <td>10</td> <td>284</td> <td>2,377</td> <td>1,703</td> <td>33,952</td> <td>22,655</td> <td>5.25</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>30</td> <td>7</td> <td>290</td> <td>2,213</td> <td>1,743</td> <td>31,452</td> <td>15,727</td> <td>4.23</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>30</td> <td>10</td> <td>283</td> <td>2,254</td> <td>1,704</td> <td>32,106</td> <td>19,913</td> <td>4.32</td> </tr> <tr> <td>&nbsp;</td> <td>pl</td> <td>20</td> <td>8</td> <td>196</td> <td>1,712</td> <td>1,181</td> <td>15,939</td> <td>13,809</td> <td>4.82</td> </tr> <tr> <td>&nbsp;</td> <td>ru</td> <td>30</td> <td>10</td> <td>294</td> <td>2,435</td> <td>1,759</td> <td>34,766</td> <td>20,509</td> <td>5.13</td> </tr> </tbody> </table> <div> <h5>Dataset statistics per speaker (average)</h5> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>16.4</td> <td>102.2</td> <td>95.3</td> <td>1,437.2</td> <td>4,047.0</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>19.8</td> <td>114.9</td> <td>114.9</td> <td>919.6</td> <td>3,945.5</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>18.5</td> <td>107.7</td> <td>107.7</td> <td>430.8</td> <td>3,658.0</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>17.4</td> <td>101.8</td> <td>101.8</td> <td>916.6</td> <td>4,870.5</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>12.5</td> <td>74.6</td> <td>74.6</td> <td>298.4</td> <td>3,141.5</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>76.3</td> <td>436.2</td> <td>436.2</td> <td>872.3</td> <td>24,892.5</td> </tr> <tr> <td>CV</td> <td>en</td> <td>9.6</td> <td>60.1</td> <td>57.8</td> <td>1,114.2</td> <td>551.5</td> </tr> <tr> <td>&nbsp;</td> <td>de</td> <td>9.5</td> <td>59.4</td> <td>57.1</td> <td>1,131.4</td> <td>568.5</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>9.7</td> <td>60.2</td> <td>58.2</td> <td>1,145.8</td> <td>618.0</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>9.9</td> <td>61.1</td> <td>59.5</td> <td>1,162.5</td> <td>612.0</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>9.5</td> <td>59.4</td> <td>56.8</td> <td>1,131.7</td> <td>610.0</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>9.7</td> <td>59.7</td> <td>58.1</td> <td>1,048.4</td> <td>409.5</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>9.4</td> <td>59.3</td> <td>56.8</td> <td>1,070.2</td> <td>564.0</td> </tr> <tr> <td>&nbsp;</td> <td>pl</td> <td>9.8</td> <td>61.1</td> <td>59.0</td> <td>797.0</td> <td>515.0</td> </tr> <tr> <td>&nbsp;</td> <td>ru</td> <td>9.8</td> <td>60.9</td> <td>58.6</td> <td>1,158.9</td> <td>392.5</td> </tr> </tbody> </table> <h2>More Information</h2> <h3>Paper</h3> <p>The paper in which this dataset is proposed will be published at Interspeech 2024.</p> <p><a href="https://arxiv.org/abs/2407.02937">The preprint is available on arXiv</a>: Meyer, Sarina, Florian Lux, and Ngoc Thang Vu. "Probing the Feasibility of Multilingual Speaker Anonymization." <em>arXiv preprint arXiv:2407.02937</em> (2024).&nbsp;</p> <h3>Code</h3> <p>All code related to this data, including the data and descriptions, as well as preparation scripts to use this data for speaker anonymization, can be found in our Github repository:<a href="https://github.com/DigitalPhonetics/speaker-anonymization/tree/multilingual"> https://github.com/DigitalPhonetics/speaker-anonymization/tree/multilingual</a></p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Fig. 2 in Rearing protocol and density trials of the brown marmorated stink bug (Hemiptera: Pentatomidae) in the laboratory

Fig. 2. Mean (± SD) number of egg masses laid per female at densities of 1:1, 2:2, and 5:5 (female to male) per cage and mean (± SD) survival of females from first egg mass laid at densities of 1:1, 2:2, and 5:5 per cage

opencc-by-4.0Dec 2016View details →
zenodo40/100

Fig. 1. Bug dorms. A in Rearing protocol and density trials of the brown marmorated stink bug (Hemiptera: Pentatomidae) in the laboratory

Fig. 1. Bug dorms. A. The large cage can house up to 12 bug dorms, with each stack containing 3 dorms. B. The large insect cage with humidifier attached. C. Stacks of Petri dishes. D. Humidifier.

opencc-by-4.0Dec 2016View details →
zenodo40/100

Figure.4. Proposed system's flow chart-Single Trial Classification of Evoked EEG Signals Due to RGB Colors

<p>In this paper we proved the possibility to perform a single trial classification the EEG signals which are evoked by the RGB color stimulus. The required time to do this process is much shorter than the time which is required by any other stimulus, such as imagery and spelling words, which is presented in the previous researches. This result proves the main idea behind using colors in the next generation of BCI systems, which is based on introducing more efficient and faster systems that are able to give a quicker response than any other time. As a future work, we are going to conduct a BCI application that controls a cursor movement on PC by using those signals. This is unlike earlier BCI systems where cursor controlled movement application is controlled by the imagination of foot and hand movement, but no one has controlled it with colored stimuli before. Such study would be used to simulate an environment where a disabled person would be expected to drive a vehicle in a virtual environment with a possible uniform background, in which the vehicle will either start and/or stop moving on appearance of Green and Red lights respectively.</p>

opencc-by-4.0Jan 2016View details →
zenodo40/100

Figure.3.Average accuracy of investigated FE methods-Single Trial Classification of Evoked EEG Signals Due to RGB Colors

<p>Each data set is recorded with 60 trails for each color from four channels, each trail contains 768 frames per channel. In order to train all the data from all channels, the trail contained 3072 frames as one vector. Then, trail by trail passed to EMD to reduce the data into a collection of intrinsic mode functions (IMF) from which the features can be extracted. Each data set represents 9 IMFs, each IMF contains lower frequency components than the previous one. In this paper, we investigate some of feature extraction methods to find out which one can give us the most reliable features. In order to know that, we trained these features with the SVM classifier and the accurate results are placed in the below tables. The classification&#39;s accurate results of the investigated feature extraction methods are shown in Figure 3. According to the accuracy of the results, we found that the best method to extract features is through the EMD residual, where the average accuracy was of 88.5% within 14 seconds. This is due to the nature of the residue as it provides the frequency representation of the delta, alpha and beta rhythms, which are the main components of ERP that respond to different color stimuli. A flow chart is inserted in Figure 4 as a summary for the used methods in this study.</p>

opencc-by-4.0Jan 2016View details →
zenodo40/100

Figure 1. Experimental protocol-Single Trial Classification of Evoked EEG Signals Due to RGB Colors

<p>Various methods exist to enhance and pre-process EEG signals by removing different artifacts like eye movement and blinking, Electrooculography (EOG) or Electromyography (EMG). The complexity of EEG signal&#39;s representation makes it difficult to define the circle that encloses most of the data points of their total. The problem with these methods is that they work at frequency domain or time domain, but not both, which causes a loss of important data during the processing stage. The researches show that the combination of frequency and time domain information can provide more completed features that improve the classification performance of EEG signals. Empirical Mode Decomposition (EMD) has recently been developed by N. (Huang Huang et al., 1998) as an adaptive time-frequency data analysis method. It has proven to be quite versatile in a broad range of applications for extracting signals from data generated in noisy nonlinear and non- stationary processes. Wavelet Transform (WT) is also an analysis method that uses the time- frequency domain. However, EMD acts essentially as a filter bank, resembling those involved in wavelet decompositions.</p>

opencc-by-4.0Jan 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record