Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
440
datasets available to search
ShareScore release 0.9.0
Dataset results
440 results for “Combined analysis”
Phlorest phylogeny derived from Birchall et al. 2016 'A combined comparative and phylogenetic analysis of the Chapacuran language family'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Birchall, Joshua, Michael Dunn, and Simon J. Greenhill. 2016. A combined comparative and phylogenetic analysis of the Chapacuran language family. International Journal of American Linguistics 82 (3): 255–84. doi: 10.1086/687383</p> </blockquote>
Data from: Combined experimental-numerical analysis of the temperature evolution and distribution during friction surfacing
<p>This dataset contains the data for the publication "Combined experimental-numerical analysis of the temperature evolution and distribution during friction surfacing".</p>
Combined unsupervised and semi-automated supervised analysis of flow cytometry data reveals cellular fingerprint associated with newly diagnosed pediatric type 1 diabetes
<p>Type 1 diabetes is a chronic autoimmune disease resulting in an immune-mediated loss of pancreatic β-cells; however, an unbiased and reproducible profiling of type 1 diabetes-specific circulating immunome at disease onset has yet to be explored. In this study, fresh whole blood was collected from a pediatric cohort of 107 patients with new-onset type 1 diabetes, 85 relatives of patients with type 1 diabetes with 0-1 islet autoantibodies, 58 patients with celiac disease or autoimmune thyroiditis and 76 healthy controls. Up to 6 mL of blood was collected from each subject into a VACUETTE® TUBE 6 ml ACD-B (Greiner). Fresh whole blood underwent red blood cell lysis, was washed and stained with specific monoclonal antibodies. Fresh whole blood samples were stained with five panels of antibodies labelled as T cells, T&NK cells, B cells, Tregs and DCs/monos encompassing main subsets of T cells, NK cells, B cells, Tregs, DCs and monocytes detected using 26 surface markers and the intracellular marker forkhead box P3 (FoxP3); for the Treg panel, intracellular staining was performed after fixation and permeabilization. Cells were acquired on a BD FACSCanto-II flow cytometer equipped with FACSDiva software (Becton Dickinson, Franklin Lakes, NJ). </p>
CLDF dataset derived from Birchall et al.'s "A Combined Comparative and Phylogenetic Analysis of the Chapacuran Language Family" from 2016
<p>Cite the source of the dataset as:</p> <blockquote> <p>Birchall J, Dunn M, & Greenhill SJ. 2016. A Combined Comparative and Phylogenetic Analysis of the Chapacuran Language Family. International Journal of American Linguistics 82(3). 255–284.</p> </blockquote>
BIKE Key-Recovery: Combining Power Consumption Analysis and Information-Set Decoding
<p>Data used in the paper: "BIKE Key-Recovery: Combining Power Consumption Analysis and Information-Set Decoding". The paper has been accepted at <a href="https://sulab-sever.u-aizu.ac.jp/ACNS2023/">ACNS-2023</a>.</p> <p>The available dataset contains a file with a few power consumption curves taken from a Cortex-M4 (STM32F4) on a CW308 board.</p> <p>The file is a numpy array stored using the np.save API.<br> The file can be directly used for running the notebooks provided in the <a href="https://github.com/benoitgerard/sca-bike">publication github</a>.</p>
Supplementary data for "Ecological assessment of combined sewer overflow management practices through the analysis of benthic and hyporheic sediment bacterial assemblages of an intermittent stream"
<p><strong>Supplementary data for the Pozzi <em>et al.</em> paper entitled "Ecological assessment of combined sewer overflow management practices through the analysis of benthic and hyporheic microbial assemblages and a tracking of exogenous bacterial taxa in a peri-urban intermittent stream".</strong></p> <p># Created by Dr Adrien C. MEYNIER POZZI on June, 29th, 2023<br> # Part of DOmic research project funded by the Agence de l’Eau - Rhône Méditerranée Corse [AE-RMC, Project 2020 0702 DOmic, 2020-2023], and of the DOmic extension funded by the EUR H2O'Lyon [ANR-17-EURE-0018] of Université de Lyon<br> # Part of the Chaudanne river long-term experiment site belonging to the Observatoire de Terrain en Hydrologie Urbaine (OTHU)<br> # Part of the work conducted in the team on Opportinistic Bacterial Pathogen in the Environment (BPOE) led by Dr. Benoit Cournoyer<br> # Samples were obtained in 2 campaigns, corresponding to periods before (2010-2011) or after (2018) the implementation of the 91/271/EEC European Directive that limited Combined-Sewer Overflow (CSO) discharges to the Chaudanne river<br> # Samples consisted in surface water, benthic and hyporheic sediments taken in run, riffle and pool geomorphologic features, either upstream or downstream the CSO outlet, plus positive and negative controls</p> <table> <tbody> <tr> <td><strong>Metadata. Name and description of data tables provided as supplementary information</strong></td> </tr> <tr> <td><strong>Data Name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>Data S1. River hydrology variables and hydraulic gradients at surveyed transects</td> <td>Array to describe the hydrologic variables and gradients at the studied transects. Top line is header, second line is metadata for each recorded variable, and third line is the unit of the variable, if any.</td> </tr> <tr> <td>Data S2. Environmental variables (water physical-chemistry, nutrients, FIBs, MTEs, PAHs) with metadata</td> <td>An array to list environmental variables for all true samples (n=90) included in the study. Sample identifiers and dates are provided. First 8 rows list the CAS number, SANDRE number, unit, method, limit of quantification and norm for each variable, if any.</td> </tr> <tr> <td>Data S3. Hydrological indices and synthetic variables computed with ClustOfVar</td> <td>Hydrological indices computed for the river flow, precipitations and CSO overflows computed over a 3-week period preceding each sampling date.</td> </tr> <tr> <td>Data S4. Discharge events selected to compute CSO dilution ratios</td> <td>An array to describe CSO events included for the computation of the CSO dilution ratio (SI Data 6A) together with 6 tables and 3 figures (SI Data 6B to 6J) describing the CSO event ratio all year round over the studied period, as well as for events that occurred before or after the CSO was modified and during low flow or high flow season. In SI Data 6A, top line is header and second line is metadata for each recorded variable.</td> </tr> <tr> <td>Data S5. Raw environmental matrix for use in R</td> <td>An array to list experimental design and environmental variables for all true samples and controls. Several environmental variables were synthetized using the ClustOfVar method (Chavent et al (2012) 10.18637/jss.v050.i13). Format is directly usable in R software.</td> </tr> </tbody> </table> <p> </p>
Combining traceological analysis and ZooMS on Early Neolithic bone artefacts from the Cave of Coro Trasito, NE Iberian Peninsula: Cervidae used equally to Caprinae
<p>MALDI-ToF MS and MALDI-tims-Q-ToF MS data (raw data: .txt files, merged spectra: msd files) used for the ZooMS analysis of 20 bone artefacts from the Early Neolithic site of Coro Trasito.</p>
Online Appendix and Cetacean Datasets for: The Occurrence Birth-Death Process for combined-evidence analysis in macroevolution and epidemiology
<p>Phylodynamic models generally aim at jointly inferring phylogenetic relationships, model parameters, and more recently, the number of lineages through time, based on molecular sequence data. In the fields of epidemiology and macroevolution these models can be used to estimate, respectively, the past number of infected individuals (prevalence) or the past number of species (paleodiversity) through time. Recent years have seen the development of "total-evidence" analyses, which combine molecular and morphological data from extant and past sampled individuals in a unified Bayesian inference framework. Even sampled individuals characterized only by their sampling time, i.e. lacking morphological and molecular data, which we call occurrences, provide invaluable information to reconstruct the past number of lineages.</p> <p>Here, we present new methodological developments around the Fossilized Birth-Death Process enabling us to (i) incorporate occurrence data in the likelihood function; (ii) consider piecewise-constant birth, death and sampling rates; and (iii) reconstruct the past number of lineages, with or without knowledge of the underlying tree. We implement our method in the RevBayes software environment, enabling its use along with a large set of models of molecular and morphological evolution, and validate the inference workflow using simulations under a wide range of conditions.</p> <p>We finally illustrate our new implementation using two empirical datasets stemming from the fields of epidemiology and macroevolution. In epidemiology, we infer the prevalence of the COVID-19 outbreak on the Diamond Princess ship, by taking into account jointly the case count record (occurrences) along with viral sequences for a fraction of infected individuals. In macroevolution, we infer the diversity trajectory of cetaceans using molecular and morphological data from extant taxa, morphological data from fossils, as well as numerous fossil occurrences. The joint modeling of occurrences and trees holds the promise to further bridge the gap between between traditional epidemiology and pathogen genomics, as well as paleontology and molecular phylogenetics.</p>
Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text
<p>Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text covering the Olympic legacy of Rio 2016 and London 2012. Data was searched via Google search engine. It is composed of sentiment labels assigned to 1271 news articles in total.</p> <p><strong>News outlets:</strong></p> <ul> <li>BBC</li> <li>Daily Mail</li> <li>The Telegraph</li> <li>The Guardian</li> <li>Globo</li> <li>Estadao</li> <li>Folha de S. Paulo</li> </ul> <p><strong>Events covered by the articles:</strong></p> <ul> <li>London 2012 Olympic legacy</li> <li>Rio 2016 Olympic legacy</li> </ul> <p>All classifiers were used in texts in English. Text originally published in Portuguese by the Brazilian media were automatically translated.</p> <p><strong>Sentiment classifiers used:</strong></p> <ul> <li>Vader</li> <li>BERT (Trained on Amazon data)</li> <li>BERT (Trained on twitter data - 140)</li> </ul> <p>Each document (spreadsheet - xlsx) refers to one outlet and one event (London 2012 or Rio 2016).</p> <p><strong>How were labels assigned to the texts?</strong></p> <p>These labels are a combination of the three sentiment classifiers listed above. If two of them agree with the same label, then this label would be considered as right. Otherwise, the label ‘other’ was assigned.</p> <p>For news article body text: the proportion of sentences of each sentiment type was used to assign labels to the whole article instead of averaging the sentence scores. For example, if the proportion of sentences with negative labels is greater than 50%, then the article is assigned a negative label.</p> <p><strong>The documents are composed of the following columns:</strong></p> <ul> <li>Rank: the position of the article on Google search ranking</li> <li>Date: date of article's publication (DD/MM/YYYY)</li> <li>Link: article's link</li> <li>Title: article's title</li> <li>Sentiment_Title: final sentiment for article headline</li> <li>Sentiment_Text: final sentiment for article's body text</li> </ul> <p><em>PS: Documents do not include articles' body text. </em></p> <p><strong>Sentiment is presented in labels as follows:</strong></p> <ul> <li>Pos: Positive</li> <li>Neg: Negative</li> <li>Neutral: Neutral</li> <li>other: inconclusive - if each of the 3 classifiers assigned a different label to the article, the label 'other' was used. Therefore, 'other' identifies contradictory results.</li> </ul> <p> </p>
Combining dynamic and static analysis for automated grading SQL statements
<p><strong>Introduction</strong></p> <p>Our experiment was conducted in an undergraduate Relational Database course at the Australian National University. The experiment was conducted on August 10th 2018 when students enrolled in the Relational Database course started to learn relational data model and SQL. The experiment was carried out fully online for three weeks and a total of 393 students were enrolled. The students were asked to login in an online assessment platform and complete 15 exercises. This platform provided an SQLite environment in students browsers by compiling the SQLite C code with Emscripten.</p> <p>Students were allowed to submit and execute their answers in the form of SQL statements. If the execution result of the statement submitted by the student is the same as that of the reference statement, the online assessment platforms will return a feedback message indicating that the execution result is correct. During the interaction with the assessment platform, statements submitted by students were recorded and archived. Overall, our experiment had collected 12,899 statements submitted by students. To create a benchmark dataset that can be used to evaluate different grading approaches, we randomly selected 45 SQL statements submitted by students for each exercise, and asked three teaching assistants to grade them manually. Finally, we average the scores provided by the three assistants and take it as the final score of each statement. The dataset collected in this experiment is ready for public release.</p> <p>All experimental data are stored in Submission.sqlite, which is an SQLite database file. It is recommended to use software such as DB browser or SQLite expert to explore the database.</p> <p> </p> <p><strong>Datatable description</strong></p> <p> </p> <p><em><strong>exercises_result</strong></em></p> <p>This datatable stores the statements submitted by students. Based on the execution result of statement, statements were divided into three categories.</p> <ul> <li>noninterpretable: the statement is non-executable.</li> <li>partially correct: the execution result of statement is different from the expected result.</li> <li>correct: the execution result of the SQL statement is the same as the expected result.</li> </ul> <p>After analyzing the correct statements, we found that the correct set contains some statements carefully constructed by students to deceive the examination system.</p> <p>Take exercise 1 as an example, the task is to answer the following questions using SQL statements.</p> <p>Question: Assume persons who were born in the same year are the same age and there is only one youngest person (with no ties/draws) in this database, who is/are the second youngest person(s) in the database? List the id(s) of the person(s).</p> <p>The reference statement to this exercise is:</p> <pre><code class="language-sql">SELECT p.id FROM person p WHERE p.year_born = (SELECT MAX(year_born) FROM person WHERE year_born < (SELECT MAX(year_born) FROM person)); </code></pre> <p>By exploring the database or trying to execute different statements, some students found that the ID of the person who met the conditions was '00000842', so the following statement was submitted.</p> <pre><code class="language-sql">select id from person where id ='00000842'; </code></pre> <p>The execution result of the above code was correct, but it was obviously not what the tutor expected. Therefore, we identified such statements as 'cheating'.</p> <p>Table 1 Description of exercises_result table.</p> <table> <thead> <tr> <th> <p><strong>field</strong></p> </th> <th> <p><strong>desc</strong></p> </th> <th> <p><strong>datatype</strong></p> </th> </tr> </thead> <tbody> <tr> <td> <p>submission_id</p> </td> <td> <p>Submission ID</p> </td> <td> <p>INT</p> </td> </tr> <tr> <td> <p>submitted_answer</p> </td> <td> <p>statement submitted by student</p> </td> <td> <p>TEXT</p> </td> </tr> <tr> <td> <p>submission_time</p> </td> <td> <p>Submission time</p> </td> <td> <p>NUM</p> </td> </tr> <tr> <td> <p>exercise_id</p> </td> <td> <p>Exercise ID</p> </td> <td> <p>INT</p> </td> </tr> <tr> <td> <p>is_correct</p> </td> <td> <p>Mark whether the statement is correct</p> </td> <td> <p>INT</p> </td> </tr> <tr> <td> <p>student_id</p> </td> <td> <p>Student ID</p> </td> <td> <p>INT</p> </td> </tr> <tr> <td> <p>category</p> </td> <td> <p>categories of statement</p> </td> <td> <p>TEXT</p> </td> </tr> </tbody> </table> <p> </p> <p><em><strong>exercises_benchmark</strong></em></p> <p>This datatable stores the scores provided by different assistants. We randomly selected 45 SQL statements submitted by students for each exercise, and asked three teaching assistants to grade them manually. Finally, we averaged the scores provided by the three assistants as the final score of each statement.</p> <p>Table 2 Description of exercises_benchmark table.</p> <table> <thead> <tr> <th> <p><strong>Field</strong></p> </th> <th> <p><strong>comment</strong></p> </th> <th> <p><strong>datatype</strong></p> </th> </tr> </thead> <tbody> <tr> <td> <p>Submission_id</p> </td> <td> <p>Submission ID</p> </td> <td> <p>INT</p> </td> </tr> <tr> <td> <p>grade</p> </td> <td> <p>grade provided by tutor</p> </td> <td> <p>REAL</p> </td> </tr> <tr> <td> <p>tutor</p> </td> <td> <p>tutor</p> </td> <td> <p>TEXT</p> </td> </tr> </tbody> </table> <p> </p> <p><em><strong>exercises_exercise</strong></em></p> <p>This datatable stores the exercises provided by tutor.</p> <p>Table 3 Description of exercises_exercise table.</p> <table> <thead> <tr> <th> <p><strong>Field</strong></p> </th> <th> <p><strong>comment</strong></p> </th> <th> <p><strong>datatype</strong></p> </th> </tr> </thead> <tbody> <tr> <td> <p>id</p> </td> <td> <p>Exercise ID</p> </td> <td> <p>INT</p> </td> </tr> <tr> <td> <p>title</p> </td> <td> <p>Title of exercise</p> </td> <td> <p>TEXT</p> </td> </tr> <tr> <td> <p>preamble</p> </td> <td> <p>Description of exercise</p> </td> <td> <p>TEXT</p> </td> </tr> <tr> <td> <p>difficulty</p> </td> <td> <p>Coefficient of difficulty</p> </td> <td> <p>integer</p> </td> </tr> <tr> <td> <p>ref</p> </td> <td> <p>Reference statement</p> </td> <td> <p>integer</p> </td> </tr> </tbody> </table> <p> </p> <p><em><strong>database schema</strong></em></p> <p>Please refer to db_schema.pdf for the database schema used in the experiment.</p> <p> </p> <p><strong>BibTex</strong></p> <p>if you want to cite our paper:</p> <p> </p> <blockquote> <pre>@article{wang2020combining, title={Combining dynamic and static analysis for automated grading SQL statements}, author={Wang, Jinshui and Zhao, Yunpeng and Tang, Zhengyi and Xing, Zhenchang}, journal={J Netw Intell}, volume={5}, number={4}, pages={179--190}, year={2020} }</pre> </blockquote>
Fig. 8 Urine miRNA profile analysis among different groups. a in A combined miRNA-piRNA signature in the serum and urine of rabbits infected with ToxoplaSMa gondii oocysts
Fig. 8 Urine miRNA profile analysis among different groups. a The volcano plot shows the individual statistically significant miRNA between acutely infected rabbits and control rabbits. In this plot, the x-axis is log2 fold-change, which shows the direction of the change (negative scale is decrease and positive scale is increase) in the levels of miRNA expression, while the y-axis is the –log10 FDR, which shows the significance of the change. b The volcano plot shows the individual statistically significant miRNA between chronically infected rabbits and control rabbits. c The volcano plot shows the individual statistically significant miRNA between acutely infected rabbits and chronically infected rabbits. d Venn diagram shows number of differentially expressed miRNA among different comparison pairs
Fig. 6 Serum miRNA profile analysis among different groups. a in A combined miRNA-piRNA signature in the serum and urine of rabbits infected with ToxoplaSMa gondii oocysts
Fig. 6 Serum miRNA profile analysis among different groups. a The volcano plot shows the individual statistically significant miRNA between acutely infected group and control group. In this plot, the x-axis is log2 fold-change, which shows the direction of the change (negative scale is decrease and positive scale is increase) in the levels of miRNA expression, while the y-axis is the –log10 FDR, which shows the significance of the change. b The volcano plot shows the individual statistically significant miRNA between chronically infected rabbits and control rabbits. c The volcano plot shows the individual statistically significant miRNA between acutely infected rabbits and chronically infected rabbits. d Venn diagram shows number of differentially expressed miRNA among different comparison pairs. FDR represents false discovery rate
Fig. 7 Serum piRNA profile analysis among different groups. a in A combined miRNA-piRNA signature in the serum and urine of rabbits infected with ToxoplaSMa gondii oocysts
Fig. 7 Serum piRNA profile analysis among different groups. a The volcano plot shows the individual statistically significant piRNA between acutely infected rabbits and control rabbits. In this plot, the x-axis is log2 fold-change, which shows the direction of the change (negative scale is decrease and positive scale is increase) in the levels of piRNA expression, while the y-axis is the –log10 FDR, which shows the significance of the change. b The volcano plot shows the individual statistically significant piRNA between chronically infected rabbits and control rabbits. c The volcano plot shows the individual statistically significant piRNA between acutely infected rabbits and chronically infected rabbits. d Venn diagram shows number of differentially expressed piRNA among different comparison pairs. FDR represents false discovery rate
Fig. 9 Urine piRNA profile analysis among different groups. a in A combined miRNA-piRNA signature in the serum and urine of rabbits infected with ToxoplaSMa gondii oocysts
Fig. 9 Urine piRNA profile analysis among different groups. a The volcano plot shows the individual statistically significant piRNA between acutely infected group and control group. In this plot, the x-axis is log2 fold-change, which shows the direction of the change (negative scale is decrease and positive scale is increase) in the levels of piRNA expression, while the y-axis is the –log10 FDR, which shows the significance of the change. b The volcano plot shows the individual statistically significant piRNA between chronically infected group and control group. c The volcano plot shows the individual statistically significant piRNA between acutely infected group and chronically infected group. d Venn diagram shows number of differentially expressed piRNA among different comparison pairs
Fig. 8 in Geometric morphometric on a new species of Trichodinidae. A tool to discriminate trichodinid species combined with traditional morphology and molecular analysis
Fig. 8. PCA. Principal component scatter plot (PCA) conducted on the elliptic Fourier descriptions of denticles shapes using the first 10 harmonics; this figure shows the first two principal components (PC1 and PC2 are on the x and y-axes, respectively).
Fig. 9 in Geometric morphometric on a new species of Trichodinidae. A tool to discriminate trichodinid species combined with traditional morphology and molecular analysis
Fig. 9. Linear discriminant analysis (LDA) of Trichodina spp. using normalized elliptical Fourier descriptors. Percentages indicate the proportion of the trace captured in each LD component.
Fig. 4. Tree derived from a in Geometric morphometric on a new species of Trichodinidae. A tool to discriminate trichodinid species combined with traditional morphology and molecular analysis
Fig. 4. Tree derived from a Maximum Likelihood (ML) analysis. The bootstrap consensus tree bases on ML inferred from 500 replicates. Bootstrap values for ML are given above nodes.
Fig. 2 in Geometric morphometric on a new species of Trichodinidae. A tool to discriminate trichodinid species combined with traditional morphology and molecular analysis
Fig. 2. Diagrammatic drawings of denticles of trichodinids. (A and B) Denticle of Trichodina bellotti n. sp. from Austrolebias bellottii. (C) Trichodina hypsilepis redrawn from Wellborn (1967). (D) Trichodina heterodentata redrawn from Duncan (1977). (E) Trichodina paraheterodentata redrawn from Tang and Zhao (2013). (F) Trichodina pseudoheterodentata redrawn from Tang et al. (2017).
Fig. 3 in Geometric morphometric on a new species of Trichodinidae. A tool to discriminate trichodinid species combined with traditional morphology and molecular analysis
Fig. 3. Phylogenetic tree based on 18S rDNA sequences by Bayesian Inference, with the model Trn + I + G applied in Mrbayes v.3.2.1. The new sequenced forms are in bold. Numbers given at nodes of branches are the posterior probability value.
Fig. 5 in Geometric morphometric on a new species of Trichodinidae. A tool to discriminate trichodinid species combined with traditional morphology and molecular analysis
Fig. 5. Denticles silhouettes utilized on Fourier analysis. Trichodina bellottii n. sp., Trichodina heterodentata redrawn from Duncan (1977); Albaladejo and Arthur, 1989; Bondad-Reantaso and Arthur, 1989; Van As and Basson, 1989; Basson and Van As, 1994; Al Rasheid et al., 2000; Asmat, 2004; Dove and O'Donoghue, 2005; Dias et al., 2009; Martins et al., 2010; Benites de Pádua et al., 2012; Miranda et al., 2012; Valladão et al., 2014. Trichodina paraheterodentata redrawn from Tang and Zhao (2013). Trichodina pseudoheterodentata redrawn from Tang et al. (2017).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.