Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,456
datasets available to search
ShareScore release 0.7.1
Dataset results
1,456 results for “parallelism”
Datasets for Parallel Meta-blocking
<p>This is a collection of the real-world datasets that were used in the publications:</p> <ul> <li>Vasilis Efthymiou, George Papadakis, George Papastefanatos, Kostas Stefanidis, Themis Palpanas: Parallel meta-blocking for scaling entity resolution over big heterogeneous data. Inf. Syst. 65: 137-157 (2017) 2015</li> <li>Vasilis Efthymiou, George Papadakis, George Papastefanatos, Kostas Stefanidis, Themis Palpanas: Parallel meta-blocking: Realizing scalable entity resolution over large, heterogeneous data. IEEE BigData 2015: 411-420</li> </ul>
Patch clamp dataset for: A massively parallel assay accurately discriminates between functionally normal and abnormal variants in a hotspot domain of KCNH2
<p><span>Many genes, including <em>KCNH2</em>, contain </span><span>'hotspot' domains associated with a high density of variants associated with disease. This has led to the suggestion that variant location can be used as evidence supporting classification of clinical variants. However, it is not known what proportion of all potential variants in hotspot domains cause loss of function. Here, we have used a massively parallel trafficking assay to characterize all single-nucleotide variants in exon 2 of <em>KCNH2, </em>a known hotspot for variants that cause long QT syndrome type 2 and an increased risk of sudden cardiac death<em>. </em>Forty-two percent of <em>KCNH2</em> exon 2 variants caused at least 50 % reduction in protein trafficking and 65% of these trafficking defective variants exerted a dominant-negative effect when co-expressed with a WT <em>KCNH2</em> allele as assessed using a calibrated patch clamp electrophysiology assay. The massively parallel trafficking assay was more accurate (AUC of 0.94) than bioinformatic prediction tools (REVEL and CardioBoost, AUC of 0.81) in discriminating between functionally normal and abnormal variants. Interestingly, over half of variants in exon 2 were found to be functionally normal, suggesting a nuanced interpretation of variants in this 'hotspot' domain is necessary. Our massively parallel trafficking assay can provide this information prospectively.</span></p>
Trafficking dataset for: A massively parallel assay accurately discriminates between functionally normal and abnormal variants in a hotspot domain of KCNH2
<p><span>Many genes, including <em>KCNH2</em>, contain </span><span>'hotspot' domains associated with a high density of variants associated with disease. This has led to the suggestion that variant location can be used as evidence supporting classification of clinical variants. However, it is not known what proportion of all potential variants in hotspot domains cause loss of function. Here, we have used a massively parallel trafficking assay to characterize all single-nucleotide variants in exon 2 of <em>KCNH2, </em>a known hotspot for variants that cause long QT syndrome type 2 and an increased risk of sudden cardiac death<em>. </em>Forty-two percent of <em>KCNH2</em> exon 2 variants caused at least 50 % reduction in protein trafficking and 65% of these trafficking defective variants exerted a dominant-negative effect when co-expressed with a WT <em>KCNH2</em> allele as assessed using a calibrated patch clamp electrophysiology assay. The massively parallel trafficking assay was more accurate (AUC of 0.94) than bioinformatic prediction tools (REVEL and CardioBoost, AUC of 0.81) in discriminating between functionally normal and abnormal variants. Interestingly, over half of variants in exon 2 were found to be functionally normal, suggesting a nuanced interpretation of variants in this 'hotspot' domain is necessary. Our massively parallel trafficking assay can provide this information prospectively.</span></p>
Underlying microevolutionary processes parallel macroevolutionary patterns in ancient Neotropical Mountains - Ecological Niche Modeling and Corridors files
<p><b>Aim</b></p> <p>Ancient climatic fluctuations are invoked as the main driving force that generates the astonishing biodiversity in ancient mountains. As a result, endemism and spatial turnover are usually high and few species are widespread among entire mountain ranges, precluding the understanding of origins of macroevolutionary patterns. Here, we used a species endemic to, but widespread in, one of the most species-rich ancient mountains on the globe to test how environmental changes acted on them and how their macroevolutionary patterns were shaped.</p> <p><b>Location</b></p> <p>Espinhaço Range, Eastern Brazil.</p> <p><b>Taxon</b></p> <p><i>Vriesea oligantha </i>species complex (Bromeliaceae).</p> <p><b>Methods</b></p> <p>We compiled data for plastidial regions and nuclear microsatellites to assess genetic diversity, population structure, migration rates and phylogenetic relationships. Using temperature and precipitation variables we modeled suitable areas for the present and the past, estimating corridors between isolated populations. We also implemented Bayesian demographic analyses to estimate ancient populations dynamics. Finally, we tested if population structure is driven by isolation by environment or by distance using a Bayesian modeling approach.</p> <p><b>Results</b></p> <p>Our results showed that the intraspecific divergence events of <i>V. oligantha</i> are older than those associated with the latest Pleistocene climatic oscillations, supporting the view that Quaternary climatic fluctuations are key components for understanding its population differentiation processes. Species distribution modeling estimated corridors between populations in the past, as also shown in the demographic analyses, depicting a major spatial reorganization during colder climates. Besides, the high genetic structure estimated results from both models of isolation by distance and by environment.</p> <p><b>Main conclusions</b></p> <p><i>V. oligantha</i> is a remarkable model to test the effects of climatic oscillations over the biological community, since this species originated in the early-Pleistocene, prevailing over several cycles of climatic fluctuations until today. The estimated demographic dynamics of <i>V. oligantha</i> agrees with the species-pump mechanism, suggesting it as the main cause of speciation within the Espinhaço Range. Moreover, the phylogeographic patterns of <i>V. oligantha</i> reflect previously recognized spatial and temporal macroevolutionary patterns in the Espinhaço Range, providing insights into how microevolutionary processes may have given rise to this astonishing mountain biodiversity.</p> <p> </p>
RF Coil Design for Accurate Parallel Imaging on 13C MRSI using 23Na Sensitivity Profiles
<p>This upload contains data for the article "RF Coil Design for Accurate Parallel Imaging on 13C MRSI using 23Na Sensitivity Profiles" published in Magnetic Resonance in Medicine, https://doi.org/10.1002/mrm.29259.</p> <p>Data are organized with respect to the figures in the article. To read and process the GE MR raw files (p-files) GE software tools are required including Matlab software from the GE MNS Research Pack. For the human data, the raw data are provided in mat-files without header information.</p> <p>All data were processed in Matlab to produce the results presented in the article. Please do not hesitate to reach out, if you are interested in any data processing methods or details, we will be happy to share relevant source code.</p>
Supplementary material for the manuscript "Simple synthesis of massively parallel RNA microarrays via enzymatic conversion from DNA microarrays"
<p>This dataset contains:</p> <ul> <li> a .txt file with the design of the Agilent SurePrint DNA microarray AMADID 086693, containing all sequences and their position on the surface </li> <li> the following raw microarray scans:</li> </ul> <ol> <li>Image of T7RNAP crystal structure, scanned at 635 nm (Cy5) (polymerase) ("01_T7RNAP image_Cy5")</li> <li>Image of T7RNAP crystal structure, scanned at 532 nm (Cy3) (dsDNA template strand) ("02_T7RNAP image_Cy3")</li> <li>Image of T7RNAP crystal structure, scanned at 488 nm (FAM) (RNA product strand) ("03_ T7RNAP image_FAM")</li> <li>Scan of the Agilent SurePrint DNA microarray (AMADID 086693) after hybridization with a Cy3-labeled oligonucleotide to untreated DNA (Block 1) and RNA as the product of the conversion process (Block 2) ("04_AgilentArray - Block 1 (untreated) vs Block 2 (converted)")</li> </ol>
Data and processing scripts for PRISM barcode sequencing data used in "Massively parallel pooled screening reveals genomic determinants of nanoparticle-cell interactions"
<p>Sequencing data for the PRISM barcodes generated after nano-particle treatment is presented in this repository alongside the code to process the sequencing counts to generate the binning probabilities and weighted scores. <br> <br> For the details please see the original publication or the bioarxiv preprint: https://doi.org/10.1101/2021.04.05.438521<br> <br> The raw data is provided in PILOT_DATA_COUNTS.csv and EXPERIMENT_DATA_COUNTS.csv files, for the pilot and the actual experiment. <br> <br> For each of these files an R script is provided to process them, along with the output of the scripts (PILOT_DATA_PROBABILITIES.csv and EXPERIMENT_DATA_PROBABILITIES.csv)</p>
Data from: Olfaction written in bone: cribriform plate size parallels olfactory receptor gene repertoires in Mammalia
The evolution of mammalian olfaction is manifested in a remarkable diversity of gene repertoires, neuroanatomy, and skull morphology across living species. Olfactory receptor genes (ORG), which initiate the conversion of odorant molecules into odor perceptions and help an animal resolve the olfactory world, range in number from a mere handful to several thousand genes across species. Within the snout, each of these ORGs is exclusively expressed by a discrete population of olfactory sensory neurons (OSN), suggesting that newly evolved ORGs may be coupled with new OSN populations in the nasal epithelium. Because OSNs axon bundles leave high-fidelity perforations (foramina) in the bone as they traverse the cribriform plate (CP) to reach the brain, we predicted that taxa with larger ORG repertoires would have proportionately expanded footprints in the CP foramina. Previous work found a correlation between ORG number and absolute CP size that disappeared when body size effects were accounted for. Using updated, digital measurement data from high-resolution CT scans and reexamining the relationship between CP and body size, we report a striking linear correlation between relative CP area and number of functional ORGs across species from all mammalian superorders. This correlation suggests strong developmental links in the olfactory pathway between genes, neurons, and skull morphology. Furthermore, because ORG number is linked to olfactory discriminatory function, this correlation supports relative CP size as a viable metric for inferring olfactory capacity across modern and extinct species. By quantifying CP area from a fossil sabertooth cat (Smilodon fatalis) we predicted a likely ORG repertoire for this extinct felid.
Richness and resilience in the Pacific: DNA metabarcoding enables parallelized evaluation of biogeographic patterns
<p><span>Islands make up a large proportion of Earth's biodiversity, yet are also some of the most sensitive systems to environmental perturbation. Biogeographic theory predicts that geologic age, area, and isolation typically drive islands' diver</span><span>sity patterns, and thus potentially impact non-native spread and community homogenization across island systems. One limitation in testing such predictions has been the difficulty of performing comprehensive inventories of island biotas and distinguishing native from introduced taxa. Here, we use DNA metabarcoding and statistical modeling as a high throughput method to survey community-wide arthropod richness, the proportion of native and non-native species, and the incursion of non-natives into primary habitats on three archipelagos in the Pacific - the Ryukyus, the Marianas and Hawaii - which vary in age, isolation and area. Diversity patterns largely match expectations based on island biogeography theory, with the oldest and most geographically connected archipelago, the Ryukyus, showing the highest taxonomic richness and the lowest proportion of introduced species. Moreover, we find evidence that forest habitats are more resilient to incursions of non-natives in the Ryukyus than in the less taxonomically rich archipelagos. Surprisingly, we do not find evidence for biotic homogenization across these three archipelagos: the assemblage of non-native species on each island is highly distinct. Our study demonstrates the potential of DNA metabarcoding to facilitate rapid estimation of biogeographic patterns, the spread of non-native species, and the resilience of ecosystems.</span></p>
ClinSpEn-CT Data: Parallel English-Spanish Biomedical Terminology
<p><strong>UPDATE August 22nd 2022: </strong>The data in this repository has been merged with the rest of the ClinSpEn data, you may access it here: https://doi.org/10.5281/zenodo.6497350</p> <p>This repository contains the sample, test and background data for the ClinSpEn-Clinical Terms sub-track. The direction of this sub-track is ES>EN.</p> <p>ClinSpEn is part of the Biomedical WMT 2022 shared task, having the aim to promote the development and evaluation of machine translation systems adapted to the medical domain with three highly relevant sub-tracks: clinical cases, medical controlled vocabularies/ontologies, and clinical terms and entities extracted from medical content.</p> <p>The terms were directly extracted from medical literature and clinical records, with particular focus on diseases, symptoms, findings, procedures and professions and translated and revised by professional medical translators.</p> <p>The sample set contains 7 000 terms as a tab-separated file (TSV), with the first column corresponding to English terms and the second column to Spanish terms.</p> <p>The test and background data is made up of a TSV file with two columns: term number and Spanish term.</p> <p>Related Links:</p> <p><strong>- Sub-track website with more information: </strong><a href="https://temu.bsc.es/clinspen/">https://temu.bsc.es/clinspen/</a></p> <p><strong>- WMT website: </strong><a href="https://www.statmt.org/wmt22/">https://www.statmt.org/wmt22/</a></p> <p><strong>- CodaLab: </strong><a href="https://codalab.lisn.upsaclay.fr/competitions/6696">https://codalab.lisn.upsaclay.fr/competitions/6696</a></p> <p> </p> <p><strong>- ClinSpEn-CC (Clinical Cases):</strong> <a href="https://doi.org/10.5281/zenodo.6497350">https://doi.org/10.5281/zenodo.6497350</a></p> <p> </p> <p><strong>- ClinSpEn-CT (Clinical Terms): </strong><a href="https://doi.org/10.5281/zenodo.6497372">https://doi.org/10.5281/zenodo.6497372</a></p> <p><strong>- ClinSpEn-OC (Ontology Concepts): </strong><a href="https://doi.org/10.5281/zenodo.6497388">https://doi.org/10.5281/zenodo.6497388</a></p>
A semi-analytical solution for heat transport in rock with parallel fractures and a heat source in both fracture and matrix
<p>In this study, we propose a two-dimensional semi-analytical solution framework based on a Green’s function approach for a flexible heat source definition, including the heat source dimensions, energy delivery strength and duration, and the presence of a heat source in the matrix and/or fracture. The solution fully accounts for heat conduction, advection, dispersion, and transient heat exchange between the fracture fluid and rock matrix in a system of parallel fractures. </p> <p>The dataset is for the figures 2-9 in the journal paper. </p> <p>The computation code is available to generate the temperature in rock matrix or fracture using Matlab. </p>
Highly Parallel Tracking Dataset
<p>The accompanying dataset for "Highly Parallel Visual Tracking: An Event-driven Particle Filter Running on SpiNNaker"<br>Arren Glover, Alan B. Stokes, Steve Furber, Chiara Bartolozzi</p> <p>the folder name corresponds to number of particles and number of events (np_nv). difficult1 corresponds to the freely moving camera dataset. Each folder contains the computed results files as well as folders for eah of the data types, event-stream (ATIS), tracking result on the cpu, and tracking result on the spinnaker. Inside each of these folders is the logged data, plus a tabular converted data. Inside the ATIS folder is also the ground truth file</p> <pre><code>@article{glover2019atis+, title={ATIS+ SpiNNaker: A fully event-based visual tracking demonstration}, author={Glover, Arren and Stokes, Alan B and Furber, Steve and Bartolozzi, Chiara}, journal={arXiv preprint arXiv:1912.01320}, year={2019} }</code></pre>
parallel striations - 14.12A
Part of a pot produced by experimental archaeologist Caroline Jeffra with the wheel-coiling technique. 3D scanning performed by archaeologist Kelly Papastergiou with the DAVID SLS-3 structured light scanner. Texture is recorded with the native DAVID-camera. The original fusion resolution is 0.0597mm and the exported OBJ is 1.01GB and consisted of 7023083 vertices. It was decimated to 438946 vertices (139MB) in order to publish the file online. Source: Objaverse 1.0 / Sketchfab
EuroparlExtract - Directional Parallel Corpora Extracted from the European Parliament Proceedings Parallel Corpus
<p>This dataset contains directional parallel corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. EuroparlExtract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>
EuroparlExtract - Comparable Corpora Extracted from the European Parliament Proceedings Parallel Corpus
<p>This dataset contains comparable translational corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. Europarl Extract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>
Real-Time Motor Unit Tracking from sEMG Signals with Adaptive ICA on a Parallel Ultra-Low Power Processor
<p>Dataset and code to replicate the paper:</p> <p>Orlandi et al., "Real-Time Motor Unit Tracking from sEMG Signals with Adaptive ICA on a Parallel Ultra-Low Power Processor"</p>
Parallel dynamics of bacterial genome reduction across independent transitions to endosymbiosis.
<p>The establishment of symbiosis dramatically alters the evolution of the associated species, making symbiotic systems ideal models for studying the impact of lifestyle changes on genomes. Here, we focused on Enterobacterales, a large and ancient bacterial lineage that includes endosymbionts with diverse host associations, ranging from gut inhabitants to intracellular environments, and from horizontal to vertical transmission. Leveraging over two hundred genomes, along with cutting-edge single-copy gene concatenation and multi-copy gene family approaches, we inferred a robust phylogenetic framework that supports eleven independent transitions to endosymbiosis. Inferences on patterns of genome evolution confirm previous hypotheses about the processes underlying genome reduction: a substantial spike in gene loss always occurs simultaneously with the establishment of endosymbiosis, while a reduction in gene acquisition mechanisms is associated with the subsequent genome erosion. Furthermore, gene family loss frequencies were correlated across independent endosymbiotic clades; genes with more conserved functions and stronger constraints on sequence evolution are lost less frequently, suggesting that differences in gene essentiality and dispensability drive the observed parallelism. Our analyses contribute to the coming of age of the theory of genome evolution in symbiotic associations and provide novel insights into the importance of recombination as an opposing force against genome erosion.</p>
Simulation results for "Localized statistics decoding: A parallel decoding algorithm for quantum low-density parity-check codes"
<p>This dataset contains simulations results presented in the paper "Localized statistics decoding: A parallel decoding algorithm for quantum low-density parity-check codes".</p> <p>The files are in `csv` file format, with data easily processable using the python library `sinter`.</p>
Figure 1 in Quantitative phosphoproteomic analysis of chicken DF-1 cells infected with Eimeria tenella, using tandem mass tag (TMT) and parallel reaction monitoring (PRM) mass spectrometry
Figure 1. Proportion of serine, threonine, and tyrosine in phosphorylation sites.
Predicting Performance and Power Consumption of Parallel Applications
<p><em><strong>Abstract: </strong>Current architectures provide many control knobs for the reduction of power consumption of applications, like reducing the number of used cores or scaling down their frequency. However, choosing the right values for these knobs in order to satisfy requirements on performance and/or power consumption is a complex task and trying all the possible combinations of these values is an unfeasible solution since it would require too much time. For this reasons, there is the need for techniques that allow an accurate estimation of the performance and power consumption of an application when a specific configuration of the control knobs values is used. Usually, this is done by executing the application with different configurations and by using these information to predict its behaviour when the values of the knobs are changed. However, since this is a time consuming process, we would like to execute the application in the fewest number of configurations possible. In this work, we consider as control knobs the number of cores used by the application and the frequency of these cores. We show that on most Parsec benchmark programs, by executing the application in 1% of the total possible configurations and by applying a multiple linear regression model we are able to achieve an average accuracy of 96% in predicting its execution time and power consumption in all the other possible knobs combinations.</em></p> <p>This dataset includes the raw data of the experiments as well as the scripts used to plot them.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.