Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,456
datasets available to search
ShareScore release 0.7.1
Dataset results
1,456 results for “parallelism”
Monthly fluorescence parallel factor analysis (PARAFAC) components for Shark River Slough, Taylor Slough, and Florida Bay, Everglades National Park (FCE LTER), Florida, USA, April 2011 - ongoing
Dissolved organic matter plays an important role in biogeochemical processes in aquatic environments such as elemental cycling, microbial loop energetics, and the transport of materials across landscapes. Since most of N (> 90%) and P (around 90%) is in the organic form in the oligotrophic subtropical Florida Coastal Everglades (FCE), study of the source and dynamics of dissolved organic matter (DOM) in the ecosystem is crucial for the better understanding of the biogeochemical cycling of nutrients. FCE are composed of estuaries with distinct regions with different biogeochemical processes. Freshwater marsh primarily receives terrestrial input and local autochthonous vegetation production. Mangrove ecotone, nevertheless, is affected by the tidal contributions from Florida Bay and local mangrove production. Florida Bay (FB) is a wedge-shaped shallow oligotrophic estuary which lays south of the Everglades, the bottom of which is covered with a dense biomass of seagrass. The sources of both freshwater and nutrients in FCE are difficult to quantify, owing to the non-point source nature of runoff from the Everglades and the dendritic cross channels in the mangroves. Furthermore, the combination of multiple DOM sources (freshwater marsh vegetation, mangroves, phytoplankton, seagrass, etc.), and the potential seasonal variability of their relative contribution, along with the history of (photo)chemical and microbial diagenetic processing, and complex advective circulation, makes the study of DOM dynamics in FCE particularly difficult using standard schemes of estuarine ecology. Quantitative information of DOM is very useful to investigate the biogeochemical cycling of DOM to a certain degree, however, qualitative information is necessary to better understand the source and dynamics of DOM. Since fluorescence spectroscopic techniques are very sensitive, quick and simple, they have been applied to investigate the fate of DOM in estuaries. Here, we have quantified a series of
Genomic evidence for the parallel regression of melatonin synthesis and signaling pathways in placental mammals
<p><strong>Supplementary Material for:</strong></p> <p>Emerling C.A., Springer M.S., Gatesy J., Jones Z., Hamilton D., Xia-Zhu D., Collin M.A., and Delsuc F. (2021). Genomic evidence for the parallel regression of melatonin synthesis and signaling pathways in placental mammals.<strong><em> Open Research Europe</em></strong> 1:75. doi:10.12688/openreseurope.13795.1.</p> <p> </p> <p><strong>Supplementary File Legends:</strong></p> <p><strong>- Supplementary_Figure_S1.pdf:</strong> <em>AANAT</em> PAML ‘master model’ showing branch categories, corresponding to “Model 1: 24 ratio” in Supplementary Table S7.</p> <p><strong>- Supplementary_Figure_S2.pdf:</strong> <em>ASMT</em> PAML ‘master model’ showing branch categories, corresponding to “Model 2: 24 ratio” in Supplementary Table S8.</p> <p><strong>- Supplementary_Figure_S3.pdf:</strong> <em>MTNR1A</em> PAML ‘master model’ showing branch categories, corresponding to “Model 1: 27 ratio” in Supplementary Table S9.</p> <p><strong>- Supplementary_Figure_S4.pdf:</strong> <em>MTNR1B</em> PAML ‘master model’ showing branch categories, corresponding to “Model 1: 46 ratio” in Supplementary Table S10.</p> <p><strong>- Supplementary_Figure_S5.pdf:</strong> RAxML <em>AANAT</em> gene tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S6.pdf: </strong>RAxML <em>ASMT</em> gene tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S7.pdf: </strong>RAxML <em>MTNR1A</em>+<em>MTNR1B</em> tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S8.pdf: </strong>Supporting data showing the inactivation of <em>MTNR1A</em> exon 2 in cetaceans. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S9.pdf: </strong>Supporting data showing the inactivation of <em>ASMT</em> in spalacids and <em>Fukomys damarensis</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S10.pdf: </strong>Supporting data showing the inactivation of <em>MTNR1A</em> in hyracoids and <em>Cyclopes didactylus</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S11.pdf: </strong>Supporting data showing the inactivation of <em>MTNR1A</em> in sirenians. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S12.pdf: </strong>Supporting data showing the inactivation of <em>AANAT</em> in sirenians and a polymorphic premature stop codon in exon 5 of <em>ASMT</em> in <em>Trichechus manatus</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S13.pdf: </strong>Supporting data showing the inactivation of <em>MTNR1A</em> in <em>Condylura cristata</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S14.pdf: </strong>Supporting data showing the inactivation of <em>MTNR1A</em> in <em>Phataginus tricuspis</em>. Read Supplementary Table S14 for further details.</p> <p><strong>- Supplementary_Figure_S15.pdf: </strong>PAML <em>AANAT</em> results, Model 1: 24 ratio (see Supplementary Table S7).</p> <p><strong>- Supplementary_Figure_S16.pdf: </strong>PAML <em>ASMT</em> results, Model 2: 24 ratio (see Supplementary Table S8).</p> <p><strong>- Supplementary_Figure_S17.pdf: </strong>PAML <em>MTNR1A</em> results, Model 1: 27 ratio (see Supplementary Table S9).</p> <p><strong>- Supplementary_Figure_S18.pdf: </strong>PAML <em>MTNR1B</em> results, Model 1: 46 ratio (see Supplementary Table S10).</p> <p><strong>- Supplementary_Table_S1.xlsx: </strong>List of species examined in this study and the sources of the genes. Source key: WGS: Sequences derived from NCBI's Whole Genome Shotgun database, with accession prefix provided; Whole Genome Sequencing of Short Reads: whole genomes were sequenced using short-read technologies. The methodologies varied for the species, and will be or have been published with other projects, so please contact the author(s) for information on the specific methodology and samples used (Xenarthrans, <em>Proteles cristatus</em>, <em>Otocyon megalotis</em>: Frédéric Delsuc, e-mail: Frederic.Delsuc@umontpellier.fr; Crocodylians: John Gatesy, e-mail: jgatesy@amnh.org; <em>Dugong dugon</em>: Mark Springer, e-mail: mark.springer@ucr.edu; SRA: sequences derived from NCBI's Sequence Read Archive; GenBank: sequences derived from NCBI's nucleotide collection; Bowhead Whale Genome Resource: sequences derived from http://www.bowhead-whale.org; Ensembl: sequences derived from Ensembl genome browser (www.ensembl.org)l; Discovar de novo: sequences derived genomes assembled via Discovar de novo (<a href="https://software.broadinstitute.org/software/discovar/blog/">https://software.broadinstitute.org/software/discovar/blog/</a>). Coverage: indicates coverage of the whole genome (reported in NCBI or other source) or individual genes (derived from short read mapping). Scaffold and contig N50: reported in NCBI or other source.</p> <p><strong>- Supplementary_Table_S2.xlsx: </strong>Accession numbers and functionality of <em>AANAT</em> in species examined. If Accession # indicated as “New”, sequence generated for this study and can be found in Supplementary Dataset S1. Parentheses after accession number indicates coordinates for sequence on the contig / scaffold. Exon colors code for the following: green = putatively functional; yellow = missing (e.g., negative BLAST results, negative mapping results); pink = one or more inactivating mutations found. Abbreviations for mutations are as follows: del = deletion; ins = insertion; start = start codon mutation; stop = premature stop codon; ? = ambiguity whether the mutation is shared among all members of the clade. Abbreviations in brackets following an inactivating mutation indicate shared inactivating mutation. Key for each abbreviation follows: Bacu = <em>Balaenoptera acutorostrata</em>; BALA = Balaenidae; BALAEN = Balaenopteridae; Bbon = <em>Balaenoptera bonaerensis</em>; CAB = <em>Cabassous</em>; Ccap = <em>Cebus capucinus</em>; CETA = Cetacea; CHLAM = Chlamyphoridae; CHOL = <em>Choloepus</em>; Cjac = <em>Callithrix jacchus</em>; CING = Cingulata; DASY = Dasypodidae; DELP = Delphinidae; DERM = Dermoptera; Erob = <em>Eschrichtius robustus</em>; INIA = <em>Inia</em>; FOLI = Folivora; GALE = <em>Galeopterus</em>; LIPO = <em>Lipotes</em>; Lobl = <em>Lagenorhynchus obliquidens</em>; MANI = Manidae; MONO = Monodontidae; MYRM = Myrmecophagidae; MYST = Mysticeti; NPP = Not present in <em>Platanista</em> or Physeteroidea, but present in other Odontocetes; NPZ = Not present in Ziphiidae, but present in other Odontocetes; Oorc = <em>Orcinus orca</em>; PEUT = Tolypeutinae; PHOC = Phocoenidae; PHOL = Pholidota; PHOR = Chlamyphorinae; PILO = Pilosa; PHYS = Physeteroidea; PONT = <em>Pontoporia</em>; Schi = <em>Sousa chinensis</em>; SIRE = Sirenia; Tadu = <em>Tursiops aduncus</em>; TOLY = <em>Tolypeutes</em>; VERM = Vermilingua; XEN = Xenarthra.</p> <p><br> <strong>- Supplementary_Table_S3.xlsx: </strong>Accession numbers and functionality of <em>ASMT</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S4.xlsx: </strong>Accession numbers and functionality of <em>MTNR1A</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S5.xlsx: </strong>Accession numbers and functionality of <em>MTNR1B</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S6.xlsx: </strong>Codon frequency model selection. These are the results from one ratio dN/dS analyses using different codon frequency models. AIC = Akaike Information Criterion.</p> <p><strong>- Supplementary_Table_S7.xlsx: </strong>Results of <em>AANAT</em> PAML dN/dS analyses for mammals. Model: BG = branch(es) grouped with background; fixed 1 = branch(es) fixed at 1. p’-value: p-value after Holm-Bonferroni correction for multiple testing. Model Comparison: if model comparison yields statistically significant differences (p < 0.05), model comparison bolded and given green background; if model comparison is still significant after Holm-Bonferroni correction, asterisk (*) added. For most models, w only shown for branch(es) of interest. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S1.</p> <p><strong>- Supplementary_Table_S8.xlsx: </strong>Results of <em>ASMT</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S2.</p> <p><strong>- Supplementary_Table_S9.xlsx: </strong>Results of <em>MTNR1A</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S3.</p> <p><strong>- Supplementary_Table_S10.xlsx: </strong>Results of <em>MTNR1B</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S4.</p> <p><strong>- Supplementary_Table_S11.xlsx: </strong>Results of PAML analyses for sauropsids.</p> <p><strong>- Supplementary_Table_S12.xlsx: </strong>Results of BLASTing and mapping short reads from <em>Alligator mississippiensis</em> RNA sequencing experiments.</p> <p><strong>- Supplementary_Table_S13.xlsx: </strong>Supporting data for validating putative inactivating mutations. Validating data came from four general sources of information: mutations shared by more than one species within a clade, mutations shared by two sources of sequencing data for the same species, mutations validated by coverage of mapped short reads and statistically elevated dN/dS ratio estimates. For additional details, see Supplementary Tables S2–S5 and S7–S10, as well as Figure 2 and Supplementary Figures S8–S18.</p> <p><strong>- Supplementary_Dataset_S1.txt:</strong><strong> </strong>Genomic alignments in fasta format used to determine the pseudogene/functional status of all four melatonin genes in different taxonomic groups.</p> <p><strong>- Supplementary_Dataset_S2.txt:</strong><strong> </strong>Alignment of <em>AANAT</em> in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML. </p> <p><strong>- Supplementary_Dataset_S3.txt: </strong>Alignment of <em>ASMT</em> in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML. </p> <p><strong>- Supplementary_Dataset_S4.txt: </strong>Alignment of <em>MTNR1A</em> and <em>MTNR1B</em> in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML. </p> <p><strong>- Supplementary_Dataset_S5.txt:</strong><strong> </strong>Codon alignments of <em>AANAT</em> used in selection pressure analyses with PAML. </p> <p><strong>- Supplementary_Dataset_S6.txt: </strong>Codon alignments of <em>ASMT</em> used in selection pressure analyses with PAML.</p> <p><strong>- Supplementary_Dataset_S7.txt:</strong><strong> </strong>Codon alignments of <em>MTNR1A</em> used in selection pressure analyses with PAML.</p> <p><strong>- Supplementary_Dataset_S8.txt: </strong>Codon alignments of <em>MTNR1B</em> used in selection pressure analyses with PAML.</p> <p><strong>- Supplementary_Dataset_S9.txt: </strong>Tree topologies in newick format used in selection pressure analyses with PAML.</p>
3D magnetotelluric modeling using high-order tetrahedral Nédélec elementson massively parallel computing platforms
<p>Accompanying data to journal article</p> <blockquote> <p>Castillo-Reyes, O., Modesto, D., Queralt, P., Marcuello, A., Ledo, J., Amor-Martin, A., de la Puente, J., García-Castillo, L.E. (2021) 3D magnetotelluric modeling using high-order tetrahedral Nédélec elements on massively parallel computing platforms. Computers & Geosciences, vol.(160): 105030 DOI: 10.1016/j.cageo.2021.105030. ISSN 0098-3004, Elsevier.</p> </blockquote>
Extended datasets from MM-IMDB and Ads-Parallelity dataset with the features from Google Cloud Vision API
<p>This is extended datasets from MM-IMDB [<a href="https://openreview.net/forum?id=S12_nquOe">Arevalo+ ICLRW'17</a>], Ads-Parallelity [<a href="https://arxiv.org/abs/1807.08205">Zhang+ BMVC'18</a>] dataset with the features from Google Cloud Vision API. These datasets are stored in jsonl (JSON Lines) format.</p> <p><strong>Abstract (from our paper):</strong></p> <p>There is increasing interest in the use of multimodal data in various web applications, such as digital advertising and e-commerce. Typical methods for extracting important information from multimodal data rely on a mid-fusion architecture that combines the feature representations from multiple encoders. However, as the number of modalities increases, several potential problems with the mid-fusion model structure arise, such as an increase in the dimensionality of the concatenated multimodal features and missing modalities. To address these problems, we propose a new concept that considers multimodal inputs as a set of sequences, namely, deep multimodal sequence sets (DM<sup>2</sup>S<sup>2</sup>). Our set-aware concept consists of three components that capture the relationships among multiple modalities: (a) a BERT-based encoder to handle the inter- and intra-order of elements in the sequences, (b) intra-modality residual attention (IntraMRA) to capture the importance of the elements in a modality, and (c) inter-modality residual attention (InterMRA) to enhance the importance of elements with modality-level granularity further. Our concept exhibits performance that is comparable to or better than the previous set-aware models. Furthermore, we demonstrate that the visualization of the learned InterMRA and IntraMRA weights can provide an interpretation of the prediction results.</p> <p><strong>Dataset (MM-IMDB and Ads-Parallelity):</strong></p> <p>We extended two multimodal datasets, namely, MM-IMDB [<a href="https://openreview.net/forum?id=S12_nquOe">Arevalo+ ICLRW'17</a>], Ads-Parallelity [<a href="https://arxiv.org/abs/1807.08205">Zhang+ BMVC'18</a>] for the empirical experiments. The MM-IMDB dataset contains 25,925 movies with multiple labels (genres). We used the original split provided in the dataset and reported the F1 scores (micro, macro, and samples) of the test set. The Ads-Parallelity dataset contains 670 images and slogans from persuasive advertisements to understand the implicit relationship (parallel and non-parallel) between these two modalities. A binary classification task is used to predict whether the text and image in the same ad convey the same message.</p> <p>We transformed the following multimodal information (i.e., visual, textual, and categorical data) into textual tokens and fed these into our proposed model. We used the <a href="https://cloud.google.com/vision">Google Cloud Vision API</a> for the visual features to obtain the following four pieces of information as tokens: (1) text from the OCR, (2) category labels from the label detection, (3) object tags from the object detection, and (4) the number of faces from the facial detection. We input the labels and object detection results as a sequence in order of confidence, as obtained from the API. We describe the visual, textual, and categorical features of each dataset below.</p> <p><em><strong>MM-IMDB</strong></em>: We used the title and plot of movies as the textual features, and the aforementioned API results based on poster images as visual features.</p> <p><em><strong>Ads-Parallelity</strong></em>: We used the same API-based visual features as in MM-IMDB. Furthermore, we used textual and categorical features consisting of textual inputs of transcriptions and messages, and categorical inputs of natural and text concrete images.</p>
Fon French Daily Dialogues Parallel Data
<p>We aim to collect, clean, and store corpora of Fon and French sentences for Natural Language Processing researches including Neural Machine Translation, Named Entity Recognition, etc. for Fon, a very low-resourced and endangered African native language.</p> <p>Fon (also called Fongbe) is an African-indigenous language spoken mostly in Benin, Togo, and Nigeria - by about 2 million people.</p> <p>As training data is crucial to the high performance of a machine learning model, the aim of this project is to compile the largest set of training corpora for the research and design of translation and NLP models involving Fon.</p> <p>Through crowdsourcing, Google Form Surveys, we gathered and cleaned #25377 parallel Fon-French# all based on daily conversations.</p> <p>To the crowdsourcing, creation, and cleaning of this version have contributed:</p> <p>1) Name: Bonaventure DOSSOU<br> Affiliation: MSc Student in Data Engineering, Jacobs University<br> Contact: femipancrace.dossou@gmail.com</p> <p>2) Name: Ricardo AHOUNVLAME<br> Affiliation: Student in Linguistics<br> Contact: tontonjars@gmail.com</p> <p>3) Name: Fabroni YOCLOUNON<br> Affiliation: Creator of the Label IamYourClounon<br> Contact: iamyourclounon@gmail.com</p> <p>4) Name: BeninLangues<br> Affiliation: BeninLangues<br> Contact: https://beninlangues.com/</p> <p>5) Name: Chris Emezue<br> Affiliation: MSc Student in Mathematics in Data Science, Technical University of Munich<br> Contact: chris.emezue@gmail.com</p> <p>_______________________________________________________</p> <p>To join as a contributor, please contact us at:<br> 1) https://twitter.com/bonadossou<br> 2) https://twitter.com/ChrisEmezue<br> 3) https://twitter.com/edAIOfficial<br> Or contact Bonaventure Dossou (femipancrace.dossou@gmail.com), Chris Emezue (chris.emezue@gmail.com)<br> _______________________________________________________</p> <p>Clavier Fongbé (WebView): https://bonaventuredossou.github.io/clavierfongbe/ (Made by Bonaventure Dossou)<br> Clavier Fongbé (Mobile Android Version): https://play.google.com/store/apps/details?id=com.fulbertodev.clavierfongbe&hl=en&gl=US (Fabroni Yoclounon, Bonventure Dossou et. al.)</p>
ENGLISH-AKUAPEM TWI PARALLEL CORPUS
<p>This dataset <em><strong>(verified_data.csv)</strong></em> is bilingual machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. <br> A transformer-based machine translator was used to generate initial translations in Akuapem Twi, which were later verified and corrected where necessary by native speakers. <br> The main idea of a typical use case for the dataset is for further training of machine translation models in Akuapem Twi.<br> The data can also be used for other downstream NLP tasks such as Named Entity Recognition and POS tagging, with appropriate additional annotations. <br> Another potential application is training unsupervised embeddings for the Akuapem Twi language.<br> In addition a higher quality 697 crowdsourced sentences <em><strong>(crowdsourced_data.csv) </strong></em>are provided for use as an evaluation set for the tasks highlighted above. It is recommended as a testing dataset for machine translation English to Twi and Twi to English models.</p> <p><strong>Acknowledgement</strong>: This project was supported by the <a href="https://www.k4all.org/project/language-dataset-fellowship/">AI4D language dataset fellowship</a> through K4all and Zindi Africa</p>
Compilation of parallel measurements comparing the temperatures recorded in Stevenson screens with those recorded in pre-Stevenson screen thermometer exposures
<p>Compilation of parallel measurements comparing the temperatures recorded in Stevenson screens with those recorded in pre-Stevenson screen thermometer exposures. This dataset accompanies Wallis et al. (2024); further details of the dataset and its creation can be found in the attached readme file and Wallis et al. (2024).</p> <p>---</p> <p><strong>References</strong></p> <p>Wallis, E.J., Osborn, T.J., Taylor, M., Jones, P.D., Joshi, M. & Hawkins, E. (2024) Quantifying exposure biases in early instrumental land surface air temperature observations. <em>International Journal of Climatology, </em>https://doi.org/10.1002/joc.8401</p>
CLDF dataset reflecting Zariquiey, Blum et al.'s "Tracing the Evolution of Panoan Languages in Parallel with Archaeological Changes in the Ucayali Basin", work in progress.
<p>Cite the source of the dataset as:</p> <blockquote> <p>Zariquiey, Roberto and Blum, Frederic and Valenzuela, Pilar and Koile, Ezequiel and Blasi, Damian and Gray, Russell and List, Johann-Mattis. "Tracing the Evolution of Panoan Languages in Parallel with Archaeological Changes in the Ucayali Basin" (work in progress).</p> </blockquote>
Real-world grasp data of a dual-arm Yumi robot with a parallel gripper and suction cup end-effectors
<p>The attached txt file contains indexes to a cleaner subset of the data issued in the first version.</p> <p>Note: <br>Version 1 contains samples with failure cases due to environment constraints, which work well for the platform used in GraspAgent 1.0 (https://doi.org/10.1109/LRA.2024.3502066). However, this can degrade the performance if used on another platform with different constraints. To solve this, version 2 reports a subset of the raw data, excluding the failure modes due to environmental causes. </p>
Highly parallel genomic selection response in replicated Drosophila melanogaster populations with reduced genetic variation
<p>Many adaptive traits are polygenic and frequently more loci contributing to the phenotype are segregating than needed to express the phenotypic optimum. Experimental evolution with replicated populations adapting to a new controlled environment provides a powerful approach to study polygenic adaptation. Since genetic redundancy often results in non-parallel selection responses among replicates, we propose a modified Evolve and Resequence (E&R) design that maximizes the similarity among replicates. Rather than starting from many founders, we only use two inbred <em>Drosophila melanogaster</em>strains and expose them to a very extreme, hot temperature environment (29°C). After 20 generations, we detect many genomic regions with a strong, highly parallel selection response in 10 evolved replicates. The X chromosome has a more pronounced selection response than the autosomes, which may be attributed to dominance effects. Furthermore, we find that the median selection coefficient for all chromosomes is higher in our two-genotype experiment than in classic E&R studies. Since two random genomes harbor sufficient variation for adaptive responses, we propose that this approach is particularly well-suited for the analysis of polygenic adaptation.</p> <p>See the README.txt file to get a description of the uploaded files. Scripts.zip contains annotated command lines and scripts for the project (see internal README.txt file).</p>
The Makerere Gendered Corpus: A Gendered English to Luganda Parallel Corpus
<p>This English-Luganda parallel sentence corpus consists of gendered examples created by a team of researchers from Makerere AI Lab at Makerere University with a team of Luganda teachers, students and freelancers. The collaborative work which involves generating English sentences under CC-0 and translating these sentences using a crowdsourcing, iterative and opensource approach was done using Pontoon an opensource Translation Management System built by Mozilla. This is a corpus of 1,000 parallel sentences.</p>
Dataset for publication "Parallel experiments in electrochemical CO2 reduction enabled by standardized analytics"
<p>Dataset for the publication: "<strong>Parallel experiments in electrochemical CO<sub>2</sub> reduction </strong><strong>enabled by standardized analytics</strong>", https://doi.org/10.1038/s41929-024-01172-x,<strong> </strong>divided by paper Figure. The dataset contains data that are both raw and processed using the open-source software available at http://dgbowl.github.io </p>
Parallel Translations from the English Wiktionary
<p>Parallel translations from English words and expressions as extracted from the English Wiktionary (database version of 2018-06-01). The data includes 2,169,063 different entries from the translation of 149,530 English words and expressions in 2,358 languages (with much variation in vocabulary size among languages: 931 languages have only one entry and German, the largest language after English, 97,091 entries). Data is offered in a tabular textual format, and all entries include (a) a unique ID, (b) a concept ID referring to the source English word, (c) a description string with the English source and a short definition (such as “dictionary/publication that explains the meanings of an ordered list of words“), (d) a language ID from the Glottolog catalog, (e) the text of the translation as given in the Wiktionary, and (f) an extra field holding complementary information, when available (such as phonetic transcription of the text, noun gender, etc.). Data is also offer in a set of files (tabular textual files, bibtex sources, and JSON metadata) following the Cross-Linguistic Data Formats (CLDF), a specification designed to allow the exchange of cross-linguistic data.</p> <p>Code for the extraction is available at http://github.com/tresoldi/wiktionary_parser and a longer description in a blog post of our group, at https://calc.hypotheses.org/?p=32</p>
MeSpEn_Parallel-Corpora
<p>MeSpEn consists of a resource of heterogeneous health related documents in Spanish and English useful to build parallel corpora for training and evaluating Spanish <-> English medical machine translation systems, to generate multilingual automatic term extraction tools, and develop other Spanish medical NLP components. MeSpEn provides the combination and harmonization of various bibliographic datasets of biomedical and clinical literature from Spain and Latin America or web-content with trusted information sources about diseases, conditions, and wellness issues for patients.</p> <p>MeSpEn was used to generate automatically bilingual health related-glossaries through automatic term detection and named entity recognition in English and target candidate term extraction in Spanish through sentence alignment approaches, implying potentially the generation of Silver Standard annotated health texts in Spanish.</p> <p>MeSpEn was used to generate automatically bilingual health related-glossaries through automatic term detection and named entity recognition in English and target candidate term extraction in Spanish through sentence alignment approaches, implying potentially the generation of Silver Standard annotated health texts in Spanish (see Villegas, et al. "The MeSpEN resource for English-Spanish medical machine translation and terminologies: census of parallel corpora, glossaries and term translations." <em>Proc. LREC 2018 Workshop MultilingualBIO: Multilingual Biomedical Text Processing)</em>.</p> <p>The MeSpEn resource aggregates several datasets, mainly from 4 principal sources: IBECS, SciELO, Pubmed and MedlinePlus:</p> <ul> <li> <p><a href="http://ibecs.isciii.es/cgi-bin/wxislind.exe/iah/online/?IsisScript=iah/iah.xis&base=IBECS&lang=i&form=F">IBECS</a> (Spanish Bibliographical Index in Health Sciences) is a bibliographical database that collects scientific journals covering multiple fields in health sciences. It is maintained by the Spanish National Health Sciences Library (BNCS), at the <a href="http://www.eng.isciii.es/">Carlos III Health Institute</a>.</p> <p>This corpus contains titles and abstracts from 168,198 records in English and Spanish. Users can find the metadata of each record written in <a href="http://dublincore.org/">Dublin Core format</a>. The original XML file of the record provided by IBECS is provided as well.</p> <p>For more information about IBECS parallel corpora, see IBECS_README file.</p> </li> <li> <p><a href="http://scielo.org/php/index.php?lang=en">SciELO</a> (Scientific Electronic Library Online) gathers electronic publications of complete full text articles from scientific journals of Latin America, South Africa and Spain. Currently is present in 15 countries and supported by the Sao Paulo Research Foundation (<a href="http://www.fapesp.br/en/">FAPESP</a>) and the Brazilian National Council for Scientific and Technological Development (<a href="http://bvsalud.org/en/">BIREME</a>).</p> <p>This corpus contains titles and abstracts from 161,710 records in English and Spanish. Users can find the metadata of each record written in <a href="http://dublincore.org/">Dublin Core format</a>.</p> <p>For more information about SciELO parallel corpora, see Scielo_README file.</p> </li> <li> <p><a href="https://www.ncbi.nlm.nih.gov/pubmed/">Pubmed</a> is a free search engine used to access the <a href="https://www.medline.com/">MedlineNLM</a>).</p> <p>This corpus contains titles and abstracts from 127,619 records. Users can find the metadata of each record written in <a href="http://dublincore.org/">Dublin Core format</a>. The original XML file of the record provided by PubMed is provided as well.</p> <p>For more information about Pubmed parallel corpora, see Pubmed_README file.</p> <p>Users can access to all Spanish articles in Pubmed by <a href="https://www.ncbi.nlm.nih.gov/pubmed?term=%22spanish%22%5BLanguage%5D">clicking here</a>. Follow these steps to download all articles' metadata in XML format:</p> <ul> <li>Click on <em>Send to</em>.</li> <li>Select <em>File</em> on <em>Choose destination</em>.</li> <li>Select <em>XML</em> on <em>Format</em>.</li> <li>And finally click on <em>Create File</em>.</li> </ul> </li> <li> <p><a href="https://medlineplus.gov/">MedlinePlus</a> is an online information service provided by the U.S. National Library of Medicine (<a href="https://www.nlm.nih.gov/ target=">NLM</a>), and gives free information about health in both English and Spanish. MedlinePlus provides the following information: <a href="https://medlineplus.gov/healthtopics.html">Health topics, </a><a href="https://medlineplus.gov/druginformation.html">Drugs and supplements, </a><a href="https://medlineplus.gov/labtests.html">Laboratory test information, </a><a href="https://medlineplus.gov/encyclopedia.html">Medical encyclopedia.</a></p> <p>There are 2 corpora available for download:</p> <ul> <li>Health topics metadata in <a href="http://dublincore.org/">Dublin Core format</a>: the source code of the site stores metadata information about each topic, we created the DC files based on these metadata. This collection contains a total of 1,063 articles in English and Spanish. For more information about it, see MedlinePlus-health-topics_README.</li> <li>Complete MedlinePlus in <a href="http://www.tei-c.org/index.xml">TEI format</a>: clean raw text and XML files of each article, structured by sections and paragraphs. This collection contains a total of 7,033 articles in English and Spanish. For more information about it, see MedlinePlus-articles_README.</li> </ul> </li> </ul> <p>These corpora are also available at http://temu.bsc.es/mespen/</p> <p>In addition, forty-six bilingual medical glossaries for various language pairs are available at https://zenodo.org/record/2205690#.XefkzdEo9hF</p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
The Makerere MT Corpus: English to Luganda parallel corpus
<p>This English-Luganda parallel sentence corpus was created by a team of researchers from AI & Data science research Lab at Makerere University with a team of Luganda teachers, students and freelancers. The collaborative work which involves generating English sentences under CC-0 and translating these sentences using a crowdsourcing, iterative and opensource approach was done using Pontoon an opensource Translation Management System built by Mozilla.</p> <p>Acknowledgment: This project was supported by the <a href="https://www.k4all.org/project/language-dataset-fellowship/">AI4D language dataset fellowship</a> through K4All and <a href="https://zindi.africa/">Zindi Africa</a>.</p>
IMPACT HTA, WP7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2 (Multi-criteria evaluation framework), Results of the 1st Web-Delphi process to HTA stakeholders, organized into 6 separate parallel panels
<p>IMPACT HTA, WP7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2 (Multi-criteria evaluation framework), Results of the 1<sup>st</sup> Web-Delphi process to HTA stakeholders, organized into 6 separate parallel panels (one panel per stakeholder group, 2 rounds), about the views of stakeholders regarding “This aspect should be considered in the evaluation of new medicines on a common basis” (2019)</p> <p>For details on the Web-Delphi process, see: IMPACT HTA, Work Package 7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2, Deliverable 7.2 (Multi-criteria evaluation framework), Advancing knowledge and MCDA tools to assist HTA agencies in evaluating medicines on a common basis (2021) Oliveira, M.D. (IST), Panos Kanavos (LSE), Bana e Costa, C. (IST)</p>
Dataset for Numerical modeling of air-vented parallel plate ionization chambers for ultra-high dose rate applications
<p>Dataset for paper: Jose Paz-Martín et al., <a href="https://www.sciencedirect.com/journal/physica-medica">Physica Medica</a> <a href="https://www.sciencedirect.com/journal/physica-medica/vol/103/suppl/C">Volume 103</a>, November 2022, Pages 147-156</p> <p><a href="https://doi.org/10.1016/j.ejmp.2022.10.006">https://doi.org/10.1016/j.ejmp.2022.10.006</a></p>
Vernacular Parallel Glosses in the Gloss-ViBe Corpus
<p>For this dataset I have collected all those instances in which there is an Old Irish gloss in the Vienna Bede that has a parallel gloss in at least one different language – Latin or Old Breton/Welsh – in one of the other three manuscripts recorded in the Gloss-ViBe corpus (https://gams.uni-graz.at/context:glossvibe). This dataset is used in <a href="https://doi.org/10.12688/openreseurope.16006.1">https://doi.org/10.12688/openreseurope.16006.1</a>.</p>
Genomic data suggest parallel dental vestigialization within the xenarthran radiation
<p><strong>Supplementary Material for:</strong></p><p>Emerling C.A., Gibb G.C., Tilak M.-K., Hughes J., Kuch M., Duggan A.T., Poinar H.N., Nachman M.W. & Delsuc F. (2023). Genomic data suggest parallel dental vestigialization within the xenarthran radiation. <i>Peer Community Journal</i> 3: e75. Recommended by <i>PCI Genomics</i>.</p><p> </p><p><strong>SCRIPTS</strong></p><p><strong>Script S1: </strong>Commands used for running CoEvol analyses on the 11 dental genes.</p><p> </p><p><strong>DATASETS</strong></p><p><strong>Dataset S1.</strong> Set of 5,262 baits used in the exon capture experiments of the 11 dental genes considered in Xenarthra.</p><p><strong>Dataset S2.</strong> ACP4 genomic alignment used for characterizing inactivating mutations.</p><p><strong>Dataset S3.</strong> AMBN genomic alignment used for characterizing inactivating mutations.</p><p><strong>Dataset S4.</strong> AMELX genomic alignment used for characterizing inactivating mutations.</p><p><strong>Dataset S5.</strong> AMTN genomic alignment used for characterizing inactivating mutations.</p><p><strong>Dataset S6.</strong> DMP1 genomic alignment used for characterizing inactivating mutations.</p><p><strong>Dataset S7.</strong> DSPP genomic alignment used for characterizing inactivating mutations.</p><p><strong>Dataset S8.</strong> ENAM genomic alignment used for characterizing inactivating mutations.</p><p><strong>Dataset S9.</strong> MEPE genomic alignment used for characterizing inactivating mutations.</p><p><strong>Dataset S10.</strong> MMP20 genomic alignments used for characterizing inactivating mutations.</p><p><strong>Dataset S11.</strong> ODAM genomic alignments used for characterizing inactivating mutations.</p><p><strong>Dataset S12.</strong> ODAPH genomic alignments used for characterizing inactivating mutations.</p><p><strong>Dataset S13.</strong> ACP4 codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S14.</strong> AMBN codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S15.</strong> AMELX codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S16.</strong> AMTN codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S17.</strong> DMP1 codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S18.</strong> DSPP codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S19.</strong> ENAM codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S20.</strong> MEPE codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S21.</strong> MMP20 codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S22.</strong> ODAM codon alignment used in Coevol and PAML analyses.</p><p><strong>Dataset S23.</strong> ODAPH codon alignment used in Coevol and PAML analyses.</p><p> </p><p><strong>SUPPLEMENTARY TABLES (Supplementary_Tables_S1-S26.xlsx)</strong></p><p><strong>Table S1.</strong> Specimen information for newly generated sequences.</p><p><strong>Table S2.</strong> Sources for DNA sequences listed for each gene, indicating methodology used. In some cases, sequences were generated via two or three methodologies. Accession numbers are associated with NCBI-derived sequences.</p><p><strong>Table S3.</strong> Primers used in PCR amplification experiments.</p><p><strong>Table S4. </strong>Results from PAML analyses (one ratio models) to determine the best codon frequency model fits. CF = codon frequency model; K = free parameters.</p><p><strong>Table S5.</strong> Inactivating mutations recorded in ACP4. The details in this caption also apply to Tables S6–S15. Taxa in bold are represented by whole genome assemblies. Exon colors code for the following: green = putatively functional; yellow = missing; pink = one or more inactivating mutations found. Abbreviations for mutations are as follows: del = deletion; ins = insertion; start = start codon mutation; stop = premature stop codon; ? = ambiguity whether the mutation is shared among all members of the clade; poly = polymorphism inferred by short reads. Abbreviations in brackets following an inactivating mutation indicate shared inactivating mutation. Key for each abbreviation follows: Bpyg = <i>Bradypus pygmaeus</i>; BRAD = <i>Bradypus</i>; Btri = <i>Bradypus tridactylus</i>; Bvar = <i>Bradypus variegatus</i>; CAB = <i>Cabassous</i>; Ccen = <i>Cabassous centralis</i>; Ccha = <i>Cabassous chacoensis</i>; CHAET = <i>Chaetophractus</i>; CHLAM = Chlamyphoridae; CHOL = <i>Choloepus</i>; Cnat = <i>Chaetophractus nationi</i>; Cuni = <i>Cabassous unicinctus</i>; Cvel = <i>Chaetophractus vellerosus</i>; Cvil = <i>Chaetophractus villosus</i>; DASY = Dasypodidae; Dkap = <i>Dasypus kappleri</i>; Dnov = <i>Dasypus novemcinctus</i>; Dpil = <i>Dasypus pilosus</i>; Dsab = <i>Dasypus sabanicola</i>; FOLI = Folivora; MYRM = Myrmecophagidae; PEUT = Tolypeutinae; PHOR = Chlamyphorinae; PHRAC = Euphractinae; PILO = Pilosa; Pmax = <i>Priodontes maximus</i>; TAM = <i>Tamandua</i>; TOLY = <i>Tolypeutes</i>; VERM = Vermilingua; XEN = Xenarthra; Zpic = <i>Zaedyus pichiy</i>.</p><p><strong>Table S6.</strong> Inactivating mutations recorded in AMBN. See additional details in Table S5 caption.</p><p><strong>Table S7.</strong> Inactivating mutations recorded in AMELX. See additional details in Table S5 caption.</p><p><strong>Table S8.</strong> Inactivating mutations recorded in AMTN. See additional details in Table S5 caption.</p><p><strong>Table S9.</strong> Inactivating mutations recorded in DMP1. See additional details in Table S5 caption.</p><p><strong>Table S10.</strong> Inactivating mutations recorded in DSPP. See additional details in Table S5 caption.</p><p><strong>Table S11.</strong> Inactivating mutations recorded in ENAM. See additional details in Table S5 caption.</p><p><strong>Table S12.</strong> Inactivating mutations recorded in MEPE. See additional details in Table S5 caption.</p><p><strong>Table S13.</strong> Inactivating mutations recorded in MMP20. See additional details in Table S5 caption.</p><p><strong>Table S14.</strong> Inactivating mutations recorded in ODAM. See additional details in Table S5 caption.</p><p><strong>Table S15.</strong> Inactivating mutations recorded in ODAPH. See additional details in Table S5 caption.</p><p><strong>Table S16.</strong> PAML results for ACP4. Model: BG = branch(es) grouped with background; fixed 1 = branch(es) fixed at 1. p-value: specific p-value only shown if lower than 0.05. Model Comparison: if model comparison yields statistically significant differences (p < 0.05), model comparison bolded and given green background. For most models, w only shown for branch(es) of interest.</p><p><strong>Table S17.</strong> PAML results for AMBN. See additional details in Table S16 caption.</p><p><strong>Table S18.</strong> PAML results for AMELX. See additional details in Table S16 caption.</p><p><strong>Table S19.</strong> PAML results for AMTN. See additional details in Table S16 caption.</p><p><strong>Table S20.</strong> PAML results for DMP1. See additional details in Table S16 caption.</p><p><strong>Table S21.</strong> PAML results for DSPP. See additional details in Table S16 caption.</p><p><strong>Table S22.</strong> PAML results for ENAM. See additional details in Table S16 caption.</p><p><strong>Table S23.</strong> PAML results for MEPE. See additional details in Table S16 caption.</p><p><strong>Table S24.</strong> PAML results for MMP20. See additional details in Table S16 caption.</p><p><strong>Table S25.</strong> PAML results for ODAM. See additional details in Table S16 caption.</p><p><strong>Table S26.</strong> PAML results for ODAPH. See additional details in Table S16 caption.</p><p> </p><p><strong>SUPPLEMENTARY FIGURES</strong></p><p><strong>Figure S1. </strong>Figure summarizing the various methodologies used to construct the various genes. Compare with Supplementary Table S2, which summarizes the methods used in the finalized constructed genes.</p><p><strong>Figure S2. </strong>Diagram showing the flow of models used in PAML dN/dS model analyses (Supplementary Tables S16–S26). This example uses real analysis results for the gene AMELX. Compare to Supplementary Table S18 and see manuscript for further details.</p><p><strong>Figure S3. </strong>Summary of the relative completeness of ACP4 for each species included in this study, within their phylogenetic context. In this figure, branch lengths are proportional to time and exon symbols are proportional to length. Bold species indicate species for which we had sequence data for this gene, and bolded branches show the nodes and branches for which we can reconstruct the gene's functional history. For each gene, all coding exons are numbered and colored according to mutation status: green = no evidence of inactivating mutations; yellow = missing data (i.e., not amplified and/or assembled); red = one or more inactivating mutations present. Note that a green exon does not necessarily mean the entire exon was recovered. Black boxes around exons of multiple species indicate the earliest shared inactivating mutations (SIMs) for a particular clade. For example, ACP4 shows exons 5 and 6 in chlamyphorids surrounded by a black box. This is due to SIMs being recorded in both exons within species in this clade, whereas there are no SIMs in common between chlamyphorids and dasypodids. To summarize the earliest examples of SIMs within a clade, we recorded the branch on which the mutational events happened. For example, in ACP4, we recorded at least 13 SIMs across five exons between <i>Dasypus novemcinctus</i> and <i>Dasypus kappleri</i>. Given that their relationships represent the basal split for all dasypodids, this suggests that ACP4 was inactivated on prior to the last common ancestor of Dasypodidae. See Supplementary Tables S5–15 and Supplementary Datasets S2-S12 for more details.</p><p><strong>Figure S4. </strong>Summary of the relative completeness of AMBN for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S5. </strong>Summary of the relative completeness of AMELX for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S6. </strong>Summary of the relative completeness of AMTN for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S7. </strong>Summary of the relative completeness of DMP1 for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S8. </strong>Summary of the relative completeness of DSPP for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S9. </strong>Summary of the relative completeness of ENAM for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S10. </strong>Summary of the relative completeness of MEPE for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S11. </strong>Summary of the relative completeness of MMP20 for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S12. </strong>Summary of the relative completeness of ODAM for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S13. </strong>Summary of the relative completeness of ODAPH for each species included in this study, within their phylogenetic context. See caption for Figure S3 for more details.</p><p><strong>Figure S14.</strong> Visualization of PAML results for ACP4. Phylogram branch lengths optimized for substitutions per codon. Red bars represent minimum dates for pseudogenization based on shared or unique inactivating mutations. Blue branches = w statistically lower than 1; blue branches with asterisk = w statistically lower than 1 and background; purple branches = statistically higher than background and lower than 1; red branches = w statistically higher than background ; red branches with asterisk = w statistically higher than background and 1; black branches that pre-date inactivating mutations = not statistically distinguishable from background or 1; black branches that post-date inactivating mutations = no statistically analyses performed on these branches.</p><p><strong>Figure S15.</strong> Visualization of PAML results for AMBN. See caption for Figure S2 for more details.</p><p><strong>Figure S16.</strong> Visualization of PAML results for AMELX. See caption for Figure S2 for more details.</p><p><strong>Figure S17.</strong> Visualization of PAML results for AMTN. See caption for Figure S2 for more details.</p><p><strong>Figure S18.</strong> Visualization of PAML results for DMP1. See caption for Figure S2 for more details.</p><p><strong>Figure S19.</strong> Visualization of PAML results for DSPP. See caption for Figure S2 for more details.</p><p><strong>Figure S20.</strong> Visualization of PAML results for ENAM. See caption for Figure S2 for more details.</p><p><strong>Figure S21.</strong> Visualization of PAML results for MEPE. See caption for Figure S2 for more details.</p><p><strong>Figure S22.</strong> Visualization of PAML results for MMP20. See caption for Figure S2 for more details.</p><p><strong>Figure S23.</strong> Visualization of PAML results for ODAM. See caption for Figure S2 for more details.</p><p><strong>Figure S24.</strong> Visualization of PAML results for ODAPH. See caption for Figure S2 for more details.</p><p><strong>Figure S25.</strong> Visualization of Coevol results for ACP4. The figure shows the Bayesian reconstruction of dN/dS across the placental phylogeny with focus on xenarthrans (armadillos, anteaters, and sloths). The variation of dN/dS was jointly reconstructed with divergence times while controlling the effect of three life-history traits (body mass, longevity, and sexual maturity). The tree is rooted with Afrotheria as the sister-group to all other placentals according to Emerling et al. (2015). Asterisks indicate non-functional sequences.</p><p><strong>Figure S26.</strong> Visualization of Coevol results for AMBN. See caption for Figure S13 for more details.</p><p><strong>Figure S27.</strong> Visualization of Coevol results for AMELX. See caption for Figure S13 for more details.</p><p><strong>Figure S28.</strong> Visualization of Coevol results for AMTN. See caption for Figure S13 for more details.</p><p><strong>Figure S29. </strong>Visualization of Coevol results for DMP1. See caption for Figure S13 for more details.</p><p><strong>Figure S30.</strong> Visualization of Coevol results for DSPP. See caption for Figure S13 for more details.</p><p><strong>Figure S31.</strong> Visualization of Coevol results for ENAM. See caption for Figure S13 for more details.</p><p><strong>Figure S32.</strong> Visualization of Coevol results for MEPE. See caption for Figure S13 for more details.</p><p><strong>Figure S33.</strong> Visualization of Coevol results for MMP20. See caption for Figure S13 for more details.</p><p><strong>Figure S34.</strong> Visualization of Coevol results for ODAM. See caption for Figure S13 for more details.</p><p><strong>Figure S35.</strong> Visualization of Coevol results for ODAPH. See caption for Figure S13 for more details.</p><p> </p>
Experiment Results: Evaluation of different parallelisation strategies with regard to the performance of parallel program execution
<p><br> Modern processors achieve an increase in performance by adding multiple cores. This means that during software development, care must be taken to parallelise the program sequences. To make predictions about the performance of a software design, there is Palladio. This is very accurate for single core processors.<br> In this bachelor thesis the influence of the chosen parallelization strategy on the performance of software is examined. Different hardware requirements are used for this purpose. They are used to generate individual work packages. These are executed by different parallelization strategies. The used parallelization strategies are: Java Threads, Java ParallelStreams, OpenMp and Akka Actor. Runtime and cache behavior are measured during each execution. In addition, the experiments are performed on different servers. The evaluation is done using acceleration curves and the Cache Miss Rate. The results show that the parallelization strategies differ only slightly in the work packages used.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.