Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
103
datasets available to search
ShareScore release 0.9.0
Dataset results
103 results for “bootstrap”
Figure 4. Bootstrap 50 in Phylogeny of Labidodemas and the Holothuriidae (Holothuroidea: Aspidochirotida) as inferred from morphology
Figure 4. Bootstrap 50% majority rule consensus tree of four trees as recovered under the equal weighting scheme. Values above branches represent bootstrap percentages (500 replicates).
Bootstrapped Lexicon of English Polarity Shifters
<p>We provide a bootstrapped lexicon of English polarity shifters and their shifting direction. We cover verbs, nouns and adjectives. Our lexicon provides 2521 shifters among a vocabulary of 9145 words, taken from WordNet v3.1 (Miller et al., 1990).</p> <p>We also provide a dataset of 2631 verb phrases that are annotated for shifting polarities. The phrases are taken from the Amazon Product Review Data corpus (Jindal & Liu, 2008).</p> <p><strong>Data</strong></p> <p><strong>1. Polarity Shifter Lexicon</strong></p> <p>A list of 9145 words, annotated for whether they are polarity shifters. Contains 2631 shifters and 6514 non-shifters.</p> <ul> <li>File: <code>shifters.txt</code></li> <li>The lexicon is a comma-separated value (CSV) table</li> <li>Each line follows the format <code>POS,LEMMA,SHIFTER_LABEL,SOURCE</code>. <ul> <li><code>POS</code>: The part of speech of the word (<code>verb</code>, <code>noun</code>, <code>adj</code>)</li> <li><code>LEMMA</code>: The lemma representation of the word in question. Multiword expressions are separated by an underscore (<code>WORD_WORD</code>).</li> <li><code>SHIFTER_LABEL</code>: Whether the word is a polarity shifter (<code>SHIFTER</code>) or a non-shifter (<code>NONSHIFTER</code>)</li> <li><code>SOURCE</code>: Whether the word was part of the gold standard (<code>GOLD_STANDARD</code>) or was bootstrapped (<code>BOOTSTRAPPED</code>). All labels, both from gold standard and bootstrap output, were verified by a human annotator.</li> </ul> </li> </ul> <p><strong>2. Sentiment Verb Phrases</strong></p> <p>A set of verb phrases, annotated for the polarity of the verb phrase and the polarity of a polar noun that it contains. Can be used to evaluate whether a polarity classifier correctly recognizes polarity shifting. The file starts with 400 phrases containing shifter verbs, followed by 2231 phrases containing non-shifter verbs.</p> <ul> <li>File: <code>sentiment_phrases.txt</code></li> <li>Every item consists of: <ul> <li>The sentence from which the VP and the polar noun were extracted.</li> <li>The VP, polar noun and the verb heading the VP.</li> <li>Constituency parse for the VP.</li> <li>Gold labels for VP and polar noun by a human annotator.</li> <li>Predicted labels for VP and polar noun by RNTN tagger (Socher et al., 2013) and <code>LEX_gold</code> approach.</li> <li>Items are separated by a line of asterisks (*)</li> </ul> </li> </ul> <p><strong>Attribution</strong><br> This dataset was created as part of the following publication:</p> <p><a href="http://marc.schulder.info/">Schulder, Marc</a> and <a href="http://www.coli.uni-saarland.de/~miwieg/">Wiegand, Michael</a> and <a href="http://ruppenhofer.de/">Ruppenhofer, Josef</a> (2020). <strong>"Automatic Generation of Lexica for Sentiment Polarity Shifters"</strong>. In: <em>Natural Language Engineering</em>. <a href="https://doi.org/10.1017/S135132492000039X">doi:10.1017/S135132492000039X</a></p> <p>If you use the data in your research or work, please cite the publication.</p>
Simulated data for paper "Conditional non-parametric bootstrap for non-linear mixed effect models"
<p>Data was simulated according to an Emax model (scenarios 1 and 2) or a Hill model (scenarios 3 and 4) with a rich (scenarios 1 and 3) and a sparse design (scenarios 2 and 4). The archive contains 4 folders with the data simulated in the first 4 scenarios (N=200 simulated datasets in each folder):<br> - scenario 1 - pdemax.rich<br> - scenario 2 - pdemax.sparse<br> - scenario 3 - pdhillhigh.rich<br> - scenario 4 - pdhillhigh.sparse<br> The data used in scenarios 5 and 6 was a subset of the datasets simulated in scenarios 3 and 4 respectively. In scenario 5, 20 subjects were taken from each dataset (subjects 1-5, 26-30, 51-55, 76-80) from the datasets in folder pdhillhigh.rich. In scenario 6, the datasets were constituted by the first 20 subjects from each sampling group of the data simulated in pdhillhigh.sparse.</p>
A new set of analytical formulae for the computation of the bootstrap current and the neoclassical conductivity in tokamaks
<p>Provided data includes data sets that have been used to develop a new set of analytical formulas for calculating the bootstrap current and the neoclassical conductivity.</p>
Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 2. - Bayesian (GTR+Γ+I and HKY+Γ models) and maximum likelihood 50% majority-rule consensus tree. Numbers in the nodes represent posterior probabilities (GTR+Γ+I and HKY+Γ, respectively), and bootstrap value for maximum likelihood and parsimony analyses, respectively. c1–Bragança, Pará; c2–Santa Maria do Pará, Pará; c3–National Forest of Amapá, Amapá; c4–Belém, Pará; i1–Solimões River, near Manaus, Amazonas; i2–Xingu River, Altamira, Pará; i3 and i4–Itacoatiara, Amazonas. MYBP–million years before present.
Figure 2. - Bayesian (GTR+Γ+I and HKY+Γ models) and maximum likelihood 50% majority-rule consensus tree. Numbers in the nodes represent posterior probabilities (GTR+Γ+I and HKY+Γ, respectively), and bootstrap value for maximum likelihood and parsimony analyses, respectively. c1–Bragança, Pará; c2–Santa Maria do Pará, Pará; c3–National Forest of Amapá, Amapá; c4–Belém, Pará; i1–Solimões River, near Manaus, Amazonas; i2–Xingu River, Altamira, Pará; i3 and i4–Itacoatiara, Amazonas. MYBP–million years before present.
Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 1. - Bayesian phylogeny of Euptychia based on one mitochondrial (COI) and one nuclear (EF1-a) gene. Posterior probabilities are listed above and bootstrap values below branches. A dash denotes bootstrap support lower than 50%. (Euptychiaattenboroughi is not included in the analysis – see text for details.)
Figure 1. - Bayesian phylogeny of Euptychia based on one mitochondrial (COI) and one nuclear (EF1-a) gene. Posterior probabilities are listed above and bootstrap values below branches. A dash denotes bootstrap support lower than 50%. (Euptychiaattenboroughi is not included in the analysis – see text for details.)
Figure 10. - Maximum-likelihood phylogeny of Epicephala species based on sequences of the COI, ArgK and EF1α genes. Numbers above nodes are maximum-likelihood bootstrap support values based on 1,000 replications. The Japanese Epicephala species are marked in blue. Symbols right to species names donate ovipositor morphology: inverted U-shape, rounded apically; inverted V-shape, acute apically.
Figure 10. - Maximum-likelihood phylogeny of Epicephala species based on sequences of the COI, ArgK and EF1α genes. Numbers above nodes are maximum-likelihood bootstrap support values based on 1,000 replications. The Japanese Epicephala species are marked in blue. Symbols right to species names donate ovipositor morphology: inverted U-shape, rounded apically; inverted V-shape, acute apically.
Bootstrap methods for quantifying the uncertainty of binding constants in the hard modeling of spectrophotometric titration data
<p>Supporting simulation code and data for the manuscript "Bootstrap methods for quantifying the uncertainty of binding constants in the hard modeling of spectrophotometric titration data"</p>
Data for the paper "Optimization of quasisymmetric stellarators with self-consistent bootstrap current and energetic particle confinement"
<p>Data for the paper "Optimization of quasisymmetric stellarators with self-consistent bootstrap current and energetic particle confinement"</p>
Supplementary Data: MPBoot: Fast phylogenetic maximum parsimony tree inference and bootstrap approximation
<p>Supplementary Data<br> MPBoot: Fast phylogenetic maximum parsimony tree inference and bootstrap approximation<br> Submitted to BMC Evolutionary Biology</p> <p>This record contains PANDIT based dataset and TreeBASE dataset (Nguyen et al. 2015) which are analyzed by different bootstrap methods in the study "MPBoot: Fast phylogenetic maximum parsimony tree inference and bootstrap approximation". The PANDIT based dataset (compressed in file data_pandit.tar.gz) is used to benchmark the accuracy of bootstrap estimates. The TreeBASE dataset (compressed in file data_treebase.tar.gz) is used to benchmark computing times and capability of finding the best-known MP scores. </p> <p>After being uncompressed, the PANDIT based dataset comprises two subdirectories corresponding to the simulated DNA and AA MSAs. They were generated by Seq-Gen (Rambaut and Grass 1997), where the model parameters and true tree were inferred from the original MSAs downloaded from the PANDIT database (Whelan et al. 2006).</p> <ul> <li>Inside "dna" subdirectory, there are 6,207 numbered directories corresponding to 6,207 DNA MSAs. Note that the numbering of these directories is not consecutive because we excluded MSAs where TNT or PAUP* runs did not finish. In each numbered directory N, there are three files: (1) data.N contains the simulated MSA in PHYLIP format; (2) model.N contains the best-fit model detected from the corresponding original MSA; (3) tree.N contains the tree (in Newick format) inferred from the corresponding original MSA. tree.N and model.N are used by Seq-Gen to simulate the MSA in data.N.</li> <li>The "aa" subdirectory is organized similarly for 6,165 AA MSAs.</li> </ul> <p>After being uncompressed, the TreeBASE dataset comprises 115 files corresponding to 115 MSAs. There are:</p> <ul> <li>70 DNA MSAs in PHYLIP format. These files follow the naming scheme dna_[number of sequences]_[number of sites].phy.</li> <li>45 protein MSAs in PHYLIP format. These files follow the naming scheme prot_[number of sequences]_[number of sites].phy.<br> </li> </ul>
Supplementary code and data for the paper `From stage to page: language independent bootstrap measures of distinctiveness in fictional speech`
<p>The repository provides full data and processing / analysis pipeline for the paper <strong>'From stage to page: language independent bootstrap measures of distinctiveness in fictional speech</strong>'<br> <br> Rendered notebooks are also available through Github:</p> <p>1) <a href="https://github.com/perechen/difs-character-voices/blob/master/data/all_stars_clean.ipynb">Preparation, energy distance and exploration</a> (main)</p> <p>2) <a href="https://github.com/perechen/difs-character-voices/blob/master/03_analysis.md">Keyword curves & formal modeling</a></p> <p> </p> <p>- `00_dracor_get_data.R`. Script uses <a href="https://dracor.org/">DraCor</a> dedicated API to get texts spoken by characters</p> <p>- `01_distinctiveness_energy.ipynb` does the heavy lifting of data wrangling, cleaning and preprocessing, plus implements energy distance bootstrapping and does exploratory analysis</p> <p>- `02_logodds_curves.R` calculates keyword curves for characters<br> <br> - `03_analysis_and_models.R` explores keyword curves and does Bayesian models</p>
Data from: Robustness of Felsenstein's versus transfer bootstrap supports with respect to taxon sampling
<p><span>The bootstrap method is based on resampling sequence alignments and re-estimating trees. Felsenstein's bootstrap proportions (FBP) is the most common approach to assess the reliability and robustness of sequence-based phylogenies. However, when increasing taxon sampling (i.e., the number of sequences) to hundreds or thousands of taxa, FBP tends to return low supports for deep branches. The Transfer Bootstrap Expectation (TBE) has been recently suggested as an alternative to FBP. TBE is measured using a continuous transfer index in [0,1] for each bootstrap tree, instead of the binary {0,1} index used in FBP to measure the presence/absence of the branch of interest. TBE has been shown to yield higher and more informative supports, while inducing a very low number of falsely supported branches.</span> <span>Nonetheless, it has been argued that TBE must be used with care due to sampling issues, especially in datasets with high number of closely related taxa. In this study, we conduct multiple experiments by varying taxon sampling and comparing FBP and TBE support values on different phylogenetic depth, using empirical datasets. Our results show that the main critique of TBE stands in extreme cases with shallow branches and highly unbalanced sampling among clades, but that TBE is still robust in most cases, while FBP is inescapably negatively impacted by high taxon sampling. We suggest guidelines and good practices in TBE (and FBP) computing and interpretation.</span></p>
Data from: Robustness of Felsenstein’s versus transfer bootstrap supports with respect to taxon sampling
Open the record for dataset details and reuse information.
Initial bootstrap of pydicom testing data
<p>An initial snapshot of the pydicom test files as found within <a href="https://github.com/SimonBiggs/pydicom/tree/b4ff70affe40fb1054b98ff049c2a6392e3e29cd/pydicom/data/test_files">https://github.com/SimonBiggs/pydicom/tree/b4ff70affe40fb1054b98ff049c2a6392e3e29cd/pydicom/data/test_files</a></p> <p>Future test files will likely be their own Zenodo submission</p>
The asymptotic behavior of bootstrap support values in molecular phylogenetics
<p>The phylogenetic bootstrap is the most commonly used method for assessing statistical confidence in estimated phylogenies by non-Bayesian methods such as maximum parsimony and maximum likelihood (ML). It is observed that bootstrap support tends to be high in large genomic datasets whether or not the inferred trees and clades are correct. Here we study the asymptotic behavior of bootstrap support for the ML tree in large datasets when the competing phylogenetic trees are equally right or equally wrong. We consider phylogenetic reconstruction as a problem of statistical model selection when the compared models are nonnested and misspecified. The bootstrap is found to have qualitatively different dynamics from Bayesian inference, and does not exhibit the polarized behavior of posterior model probabilities, consistent with the empirical observation that the bootstrap is more conservative than Bayesian probabilities. Nevertheless bootstrap support similarly shows fluctuations among large datasets, with no convergence to a point value, when the compared models are equally right or equally wrong. Thus in large datasets strong support for wrong trees or models is likely to occur. Our analysis provides a partial explanation for the high bootstrap support values for incorrect clades observed in empirical data analysis.</p>
FIGURE 58. Maximum parsimony strict consensus tree. Bootstrap proportions from 1000 in Taxonomic revision and phylogeny of the ant genus Prenolepis (Hymenoptera: Formicidae)
FIGURE 58. Maximum parsimony strict consensus tree. Bootstrap proportions from 1000 replicates are presented along the branches.
FIGURE 1. Maximum Likelihood tree with bootstrap values over 50 in Onthophagus (Palaeonthophagus) medius (Kugelann, 1792) — a good western palaearctic species in the Onthophagus vacca complex (Coleoptera: Scarabaeidae: Scarabaeinae: Onthophagini)
FIGURE 1. Maximum Likelihood tree with bootstrap values over 50% shown above the branches that resulted from the analysis of the mitochondrial DNA sequences. Collecting sites: Bulgaria: L1—Sakar Mts: Topolovgrad, L2—Strandza Mts.: Zvezdec, L3—Bakadzicite: Vojnika, L4—Sinemorec (8 km S Ahtopol); Germany: L5—Klein Schmölen, L6— Mecklenburg: Sternberg; Italy: L7—Torino (Ipla), L8—Sardinia: Rio Antas: Flumini-magg.-Tempio di Antas, L9—St. Antíoco, Tonnara, L10—Cala Gonone, L11—Monte Lupone, L12—Rocca Massima, L13—Lago Sefro, L14—Piane della Regna, L15—Colle dell'Orso, L16—Campodimele env., L17—Pantano della Zottola, Iserina; Spain: L18— Higuera de la Sierra, L19—Avila to El Barraco, L20—Guadix, L21—Extremadura: 1 km E Jarandilla.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.