Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,310
datasets available to search
ShareScore release 0.9.0
Dataset results
1,310 results for “construction”
Snapgene files of the genetic constructions for Morella thermoacetica used in Task 2.2 (Advanced GMO preparation for gas (CO2/CO/H2) fermentation process)
<p>Part of WP2 these snapgene files of the genetic constructions for Morella thermoacetica were used in Task 2.2 (Advanced GMO preparation for gas (CO2/CO/H2) fermentation process). The strains we were able to construct were:</p> <p>1.- M. thermoacetica DSM 521 pK18ACSACassette complete<br>2.- M. thermoacetica DSM 521 pK18ACSACassette AatA<br>3.- M. thermoacetica DSM 2955 pK18ACSACassette PudL-AckA<br>4.- M. thermoacetica DSM 2955 pK18ACSACassette AatA</p> <p>To construct these strains CSIC used a modular synthetic construction carrying the regions up and down of the acsA gene to delete it. Within these regions, a gene cluster composed by the GmR, pduL, ackA and aatA genes, all of them controlled by the strong constitutional promoter of the glyceraldehyde 3-phosphate dehydrogenase gene from Moorella were inserted. The modular design of our recombinant cassette allows us to construct different cassettes with different gene combinations. In all cases, the acsA gene will be deleted and the acetate cannot be transformed back into acetyl-CoA:</p> <p>Cassette 1.- ACSA-GmR-pduL-ackA-aatA-ACSA: this is a complete cassette which carries the genes encoding the phosphotransacetylase, the acetate kinase and the acetate exporter. This strain should accumulate a higher concentration of acetate inside the cell that should be exported to the medium.</p> <p>Cassette 2.- ACSA-GmR-pduL-ackA-ACSA: this cassette lacks the transporter, so this strain should accumulate a higher concentration inside the cell. Then, the only way to prevent acetate toxicity is by increasing its own mechanisms to export this compound or by growing slower in order to manage acetate accumulation.</p> <p>Cassette 3.- ACSA-GmR-aatA-ACSA: this cassette only carries the acetate exporter, so virtually all the acetate produced inside the cell should be exported. This is another strategy to increase the metabolic flux inside the cell towards acetate without overexpressing any other gene of the acetate pathway, and at the same time avoiding toxicity problems.</p> <p>Cassette 4.- ACSA-GmR-ACSA: this cassette would generate an insertion mutant in which acsA is only replaced by the GmR encoding gene. This mutant should also generate a higher amount of acetate since it cannot be transformed back to acetyl-CoA.</p>
Data for Project 'Diagnostic Accuracy, Reliability, and Construct Validity of the German Quick Mild Cognitive Impairment Screen'
<p>Data for Project 'Diagnostic Accuracy, Reliability, and Construct Validity of the German Quick Mild Cognitive Impairment Screen' consisting of (1) the complete data set of all data analyzed for the project 'Diagnostic Accuracy, Reliability, and Construct Validity of the German Quick Mild Cognitive Impairment Screen' ('Data_Brain-IT-Validation-Qmci_for-publication.xlsx'; and (2) a corresponding README file including (a) general information, (b) data and file overview, (c) sharing and access information, (d) methodological information, and (e) data-specific information.</p>
Dataset for the paper "A framework for robotic excavation and dry stone construction using on-site materials"
<p>Stone data from the <i>Science Robotics</i> paper "A framework for robotic excavation and dry stone construction using on-site materials" containing:</p><ul><li>Mesh files of 1,100 stones (quarried boulders, erratics, and concrete debris) that were digitized by the autonomous excavator HEAP<ul><li><a href="https://zenodo.org/api/records/10038881/draft/files/1100%20Unprocessed%20Stone%20Meshes.zip/content">1100 Unprocessed Stone Meshes.zip: </a>Raw mesh files directly from the poisson reconstruction of accumulated LiDAR points, containing some artifacts and floating geometries</li><li><a href="https://zenodo.org/api/records/10038881/draft/files/1100%20Closed%20Stone%20Meshes.zip/content">1100 Closed Stone Meshes.zip: </a>Clean, closed, downsampled meshes</li><li><a href="https://zenodo.org/api/records/10038881/draft/files/Stone_Shape_Properties.csv/content">Stone_Shape_Properties.csv: </a>Properties file with a list of the stone IDs (IDs in the 1xxx and 3xxx range typically correspond to concrete elements) and select shape properties</li></ul></li><li>A dataset of candidate placements from automatically generated stone walls. The candidate placement data zip files contain:<ul><li><a href="https://zenodo.org/api/records/10038881/draft/files/Candidate_Placement_Data-npy.zip/content">Candidate_Placement_Data-npy.zip: </a>SDF (.npy) representation of each candidate, with three channels of 32x32x32 for distances to the stone, the already-placed stones, and the target wall</li><li><a href="https://zenodo.org/api/records/10038881/draft/files/Candidate_Placement_Data-pcd.zip/content">Candidate_Placement_Data-pcd.zip: </a>Point cloud (.pcd) representations of each candidate, with separate files for the placed stone, target wall (search volume), and already-placed stones (where they exist)</li></ul></li><li><a href="https://zenodo.org/api/records/10038881/draft/files/sdf_classifier.zip/content">sdf_classifier.zip: </a>Python examples:<ul><li>Rendering the three channel SDF data to mesh geometry using marching cubes and libigl</li><li>Candidate SDF classification using the pretrained model</li></ul></li><li>Candidate attributes and labels<ul><li><a href="https://zenodo.org/api/records/10038881/draft/files/candidate_attributes_labels.csv/content">candidate_attributes_labels.csv: </a>CSV file containing a list of UUID's corresponding to each candidate placement in the dataset. For each candidate, additional information is included about the dimensions and location of the solution, together with the (subjectively) hand-labelled binary value for placement viability. </li></ul></li><li><a href="https://zenodo.org/api/records/10038881/draft/files/README.md/content">README.md: </a>An additional readme with some details on the attributes file</li></ul><p>If you use this data in your research, please cite the <a href="https://www.science.org/doi/10.1126/scirobotics.abp9758">journal article</a>.</p><p> </p>
Replication package for the paper: "A Study on the Pythonic Functional Constructs' Understandability"
<h1>Replication Package for "<em>A Study on the Pythonic Functional Constructs' Understandability</em>" to appear at ICSE 2024</h1> <ul> <li><strong>Authors</strong>: Cyrine Zid, Fiorella Zampetti, Giuliano Antoniol, Massimiliano Di penta</li> <li><strong>Article Preprint:</strong> <a href="https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf">https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf</a></li> <li><strong>Artifacts:</strong> <a href="https://doi.org/10.5281/zenodo.8191782">https://doi.org/10.5281/zenodo.8191782</a></li> <li><strong>License</strong>: <a>GPL V3.0</a></li> </ul> <p>This package contains folders and files with code and data used in the study described in the paper. In the following, we first provide all fields required for the submission, and then report a detailed description of all repository folders.</p> <h2>Artifact Description</h2> <h3>Purpose</h3> <p>The artifact is about a controlled experiment aimed at investigating the extent to which Pythonic functional constructs have an impact on source code understandability. The artifact archive contains:</p> <ol> <li>The material to allow replicating the study (see Section <a>Experimental-Material</a>)</li> <li>Raw quantitative results, working datasets, and scripts to replicate the statistical analyses reported in the paper. Specifically, the executable part of the replication package reproduces figures and tables of the quantitative analysis (RQ1 and RQ2) of the paper starting from the working datasets.</li> <li>Spreadsheets used for the qualitative analysis (RQ3).</li> </ol> <p>We apply for the following badges:</p> <ul> <li><strong>Available and reusable</strong>: because we provide all the material that can be used to replicate the experiment, but also to perform the statistical analyses and the qualitative analyses (spreadsheets, in this case)</li> </ul> <h3>Provenance</h3> <ul> <li><strong>Paper preprint link:</strong> <a href="https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf">https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf</a></li> <li><strong>Artifacts:</strong> <a href="https://doi.org/10.5281/zenodo.8191782">https://doi.org/10.5281/zenodo.8191782</a></li> </ul> <h3>Data</h3> <p>Results have been obtained by conducting the controlled experiment involving <a href="https://www.prolific.com/">Prolific</a>workers as participants. Data collection and processing followed a protocol approved by the University ethical board. Note that all data enclosed in the artifact is completely anonymized and does not contain sensible information.</p> <p>Further details about the provided dataset can be found in the Section <a>Results' directory and files</a></p> <h3>Setup and Usage (for executable artifacts):</h3> <p>See the Section <a>Scripts to reproduce the results, and instructions for running them</a></p> <h2><a>Experiment-Material/</a></h2> <p>Contains the material used for the experiment, and, specifically, the following subdirectories:</p> <h3><a>Google-Forms/</a></h3> <p>Contains (as PDF documents) the questionnaires submitted to the ten experimental groups.</p> <h3><a>Task-Sources/</a></h3> <p>Contains, for each experimental group (G-1...G-10), the sources used to produce the Google Forms, and, specifically: - The cover letter (Letter.docx). - A directory for each experimental task (Lambda 1, Lambda 2, Comp 1, Comp 2, MRF 1, MRF 2, Lambda Comparison, Comp Comparison, MRF Comparison). Each directory contains: (i) the exercise text (in both Word and .txt format), the source code snippet, and its .png image to be used in the form. <strong>Note:</strong> the "Comparison" tasks do not have any exercise as the purpose is always the same, i.e., to compare the (perceived) understandability of the snippets and return the results of the comparison.</p> <h3><a>Code-Examples-Table1/</a></h3> <p>Contains the source code snippets used as objects of the study (the same you can find under "Task-Sources/"), named as reported in Table 1.</p> <h2>Results' directory and files</h2> <p> </p> <h3><a>raw-responses/</a></h3> <p>Contains, as spreadsheets, the raw responses provided by the study participants through Google forms.</p> <h3><a>raw-results-RQ1/</a></h3> <p>Contains the raw results for RQ1. Specifically, the directory contains a subdirectory for each group (G1-G10). Each subdirectory contains: - For each user (named using their Prolific IDs, a directory containing, for each question (Q1-Q6) the produced python code (Qn.py) its output (QnR.txt) and its StdErr output (QnErr.txt). - "expected-outputs/": A directory containing the expected outputs for each task (Qn.txt).</p> <h4><a>working-results/RQ1-RQ2-files-for-statistical-analysis/</a></h4> <p>Contains three .csv files used as input for conducting the statistical analysis and drawing the graphs for addressing the first two research questions of the study. Specifically:</p> <ul> <li> <p><a>ConstructUsage.csv</a> contains the declared frequency usage of the three functional constructs object of the study. This file is used to draw Figure 4. The file contains an entry for each participant, reporting the (text-coded) frequency of construct usage for Comprehension, Lambda, and MRF.</p> </li> <li> <p><a>RQ1.csv</a> contains the collected data used for the mixed-effect logistic regression relating the use of functional constructs with the correctness of the change task, as well as the logistic regression relating the use of map/reduce/filter functions with the correctness of the change task. The csv file contains an entry for each answer provided by each subject, and features the following columns:</p> <ul> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>User</em>: user ID</li> <li><em>Time</em>: task time in seconds</li> <li><em>Approvals</em>: number of approvals on previous tasks performed on Prolific</li> <li><em>Student</em>: whether the participant declared themselves as a student</li> <li><em>Section</em>: section of the questionnaire (lambda, comp, or mrf)</li> <li><em>Construct</em>: specific construct being presented (same as "Section" for lambda and comp, for mrf it says whether it is a map, reduce, or filter)</li> <li><em>Question</em>: question id, from Q1 to Q6, indicate the ordering of the question</li> <li><em>MainFactor</em>: main factor treatment for the given question - "f" for functional, "p" for procedural counterpart</li> <li><em>Outcome</em>: TRUE if the task was correctly performed, FALSE otherwise</li> <li><em>Complexity</em>: cyclomatic complexity of the construct (empty for mrf)</li> <li><em>UsageFrequency</em>: usage frequency of the given construct</li> </ul> </li> <li> <p><a>RQ1Paired-RQ2.csv</a> contains the collected data used for the ordinal logistic regression of the relationship between the perceived ease of understanding of the functional constructs and (i) participants' usage frequency, and (ii) constructs' complexity (except for map/reduce/filter). The file features a row for each participant, and the columns are the following:</p> <ul> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>User</em>: user ID</li> <li><em>Time</em>: task time in seconds</li> <li><em>Approvals</em>: number of approvals on previous tasks performed on Prolific</li> <li><em>Student</em>: whether the participant declared themselves as a student</li> <li><em>LambdaF</em>: result for the change task related to a lambda construct</li> <li><em>LambdaP</em>: result for the change task related to the procedural counterpart of a lambda construct</li> <li><em>CompF</em>: result for the change task related to a comprehension construct</li> <li><em>CompP</em>: result for the change task related to the procedural counterpart of a comprehension construct</li> <li><em>MrfF</em>: result for the change task related to an MRF construct</li> <li><em>MrfP</em>: result for the change task related to the procedural counterpart of a MRF construct</li> <li><em>LambdaComp</em>: perceived understandability level for the comparison task (RQ2) between a lambda and its procedural counterpart</li> <li><em>CompComp</em>: perceived understandability level for the comparison task (RQ2) between a comprehension and its procedural counterpart</li> <li><em>MrfComp</em>: perceived understandability level for the comparison task (RQ2) between a MRF and its procedural counterpart</li> <li><em>LambdaCompCplx</em>: cyclomatic complexity of the lambda construct involved in the comparison task (RQ2)</li> <li><em>CompCompCplx</em>: cyclomatic complexity of the comprehension construct involved in the comparison task (RQ2)</li> <li><em>MrfCompType</em>: type of MRF construct (map, reduce, or filter) used in the comparison task (RQ2)</li> <li><em>LambdaUsageFrequency</em>: self-declared usage frequency on lambda constructs</li> <li><em>CompUsageFrequency</em>: self-declared usage frequency on comprehension constructs</li> <li><em>MrfUsageFrequency</em>: self-declared usage frequency on MRF constructs</li> <li><em>LambdaComparisonAssessment</em>: outcome of the manual assessment of the answer to the "check question" required for the lambda comparison ("yes" means valid, "no" means wrong, "moderate<em>chatgpt" and "extreme</em>chatgpt" are the results of GPTZero)</li> <li><em>CompComparisonAssessment</em>: as above, but for comprehension</li> <li><em>MrfComparisonAssessment</em>: as above, but for MRF</li> </ul> </li> </ul> <h3><a>working-results/inter-rater-RQ3-files/</a></h3> <p>This directory contains four .csv files used as input for computing the inter-rater agreement for the manual labeling used for addressing RQ3. Specifically, you will find one file for each functional construct, i.e., comprehension.csv, lambda.csv, and mrf.csv, and a different file used for highlighting the reasons why participants prefer to use the procedural paradigm, i.e., procedural.csv.</p> <h3><a>working-results/RQ2ManualValidation.csv</a></h3> <p>This file contains the results of the manual validation being done to sanitize the answers provided by our participants used for addressing RQ2. Specifically, we coded the behaviour description using four different levels: (i) correct ("yes"), (ii) somewhat correct ("partial"), (iii) wrong ("no"), and (iv) automatically generated. The file features a row for each participant, and the columns are the following:</p> <ul> <li><em>ID</em>: ID we used to refer the participant in the paper's qualitative analysis</li> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>ProlificID</em>: user ID</li> <li><em>Comparison for lambda construct description</em>: answer provided by the user for the lambda comparison task</li> <li><em>Final Classification</em>: our assessment of the lambda comparison answer</li> <li><em>Comparison for comprehension description</em>: answer provided by the user for the comprehension comparison task</li> <li><em>Final Classification</em>: our assessment of the comprehension comparison answer</li> <li><em>Comparison for MRF description</em>: answer provided by the user for the MRF comparison task</li> <li><em>Final Classification</em>: our assessment of the MRF comparison answer</li> </ul> <h3><a>working-results/RQ3ManualValidation.xlsx</a></h3> <p>This file contains the results of the open coding applied to address our third research question. Specifically, you will find four sheets, one for each functional construct and one for the procedural paradigm. Each sheet reports the provided answers together with the categories assigned to them. Each sheet contains the following columns:</p> <ul> <li><em>ID</em>: ID we used to refer the participant in the paper's qualitative analysis</li> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>ProlificID</em>: user ID (as in the tables from the quantitative analysis)</li> <li>: question asked to the user</li> <li><em>Final Classification</em>: The outcome of our categorization according to the taxonomy shown in Table 10.</li> </ul> <h2>Scripts to reproduce the results and instructions for running them</h2> <p> </p> <h3><a>FuncConstructs-Statistics.r</a></h3> <p>This file contains an R script that you can reuse to re-run all the analyses conducted and discussed in the paper.</p> <h3><a>FuncConstructs-Statistics.ipynb</a></h3> <p>This file contains the code to re-execute all the analysis conducted in the paper as a Jupyter Notebook (using the R Kernel).</p> <h3><a>run-analysis.sh</a></h3> <p>This script can be used to run the R script <code>FuncConstructs-Statistics.r</code> using a Docker container (see Option 1 below) in Unix operating systems.</p> <h3><a>run-analysis.bat</a></h3> <p>This script can be used to run the R script <code>FuncConstructs-Statistics.r</code> using a Docker container (see Option 1 below) in Windows operating systems (power shell recommended).</p> <h3><a>run-jupyter-container.sh</a></h3> <p>This script can be used to run a local Jupyter server (with R kernel and all required packages) from a Docker container (see Option 3 below) in Unix operating system.</p> <h3><a>run-jupyter-container.bat</a></h3> <p>This script can be used to run a local Jupyter server (with R kernel and all required packages) from a Docker container (see Option 3 below) in Windows operating systems (power shell recommended).</p> <h3>How to Run the scripts</h3> <p>There are four options to run the scripts. In all cases, one has first to open a shell terminal window (e.g., bash or sh in Unixes) in the replication package directory. For Windows, we suggest to use a Power Shell.</p> <ol> <li> <p><strong>Running the R script using Dockerized R installation</strong>: this is the simplest option, and it simply requires a running Docker engine. In <strong>Unix (MacOS, Linux)</strong>, to produce the results, one has to run the shell script "run-analysis.sh" (e.g., by typing <code>sh run-analysis.sh</code> or simply <code>./run-analysis.sh</code> after making it executable). In <strong>Windows</strong>, one has to run the script "run-analysis.bat" instead (by typing <code>.\run-analysis.bat</code>). This script (either .sh or .bat) will:</p> <ul> <li>Pull a docker image named <code>mdipenta/rexp</code> which contains an R installation with all required packages.</li> <li> <p>Run R from the container created from the image and produce the paper's results under a directory named <code>results/</code>.</p> </li> <li> <p><strong>Note:</strong> An alternative would be to run everything from inside the container, after running it in interactive mode. To this aim, please execute the following commands:</p> <ol> <li>In <strong>Unix</strong>: <code>docker run -v${PWD}:/data --rm -ti --name shell mdipenta/rexp:latest bash</code>in <strong>Windows</strong>: <code>docker run -v %cd%:/data --rm -ti --name shell mdipenta/rexp:latest bash</code></li> <li><code>cd data</code></li> <li><code>R --no-save < FuncConstructs-Statistics.r</code> After exiting the container, the "results" directory will be again populated with the study results.</li> </ol> </li> </ul> </li> <li> <p><strong>Running the R script from own R installation</strong>: this option works if one has an R installation already (or wants to use an R installation) without relying on the Docker image. The steps to be followed are:</p> <ul> <li>Uncomment the <code>install.packages(..)</code> instruction in the first lines of the script. This will allow for the installation of the required packages.</li> <li>Just run, from the current directory, the script <code>FuncConstructs-Statistics.r</code>, using the command <code>Rscript FuncConstructs-Statistics.r</code> (making sure the directory containing Rscript is in your PATH, this should work fine in Unixes, it might require to modify the PATH environment variable in Windows). Should you experience problems with the first part of the script (installations), try to execute the <code>install.packages(..)</code> statement from your R GUI, and then run the script again.</li> </ul> </li> <li> <p><strong>Using the Jupyter Notebook using a Dockerized Jupyter lab with R kernel</strong>: this option allows for opening the Jupyter Notebook with all results without having to install Jupyter with the R kernel, nor all the required R packages. The steps required are:</p> <ul> <li>Run the <code>./run-jupyter-container.sh</code> (<strong>Unix</strong>) or <code>.\run-jupyter-container.bat</code> (<strong>Windows</strong>). It will download the <code>mdipenta/myjupyter</code> image and run Jupyter lab from it.</li> <li>Open a browser on <a href="http://localhost:8888/">localhost:8888</a> (or if it does not work, <a href="http://127.0.0.1:8888/">127.0.0.1:8888)</a> and, when being asked for a password, type <code>docker</code>.</li> <li>From the Jupyter lab page, open the "FuncConstruct-Statistics.ipynb" notebook, and (if you wish) re-run it, or simply browse its results. Note: differently from options 1 and 2, results are not saved, but just displayed in the notebook.</li> </ul> </li> <li> <p><strong>Using the Jupyter Notebook from your installation</strong>: this is similar to Option 3, but it can work if you have already Jupyter lab installed, with the R kernel enabled (for details see: <a href="https://github.com/IRkernel/IRkernel">https://github.com/IRkernel/IRkernel</a>). The steps to follow are:</p> <ul> <li>Run jupyter lab (e.g., jupyter lab from the command line) and open it on a webpage.</li> <li>Open the <code>FuncConstruct-Statistics.ipynb</code> notebook.</li> <li>If you want to re-execute it, uncomment the <code>install.packages()</code> line.</li> <li>Re-run it (if you wish).</li> </ul> </li> </ol> <h3>The output</h3> <p>If using Option 1 or 2, the <code>results</code> directory will contain the following files:</p> <ul> <li><strong>Figures 4 and 5</strong> as in the paper.</li> <li><strong>Tables 2-9</strong> as in the paper in various formats (csv, tex, and for Tables 3-5 also .txt). Some notes: The diagnostics (top part, up to "Fixed effects") for Tables 3-5 are shown in the .txt files only. However, these files do not report the "OR" columns that correspond to exp(Estimate). This is because the .txt file contains the statistics dump which does not include the ORs. The .csv and .tex tables report the Fixed effects as shown in the paper, including the ORs.</li> <li><strong>rq1-rq2-correlation</strong> (.tex and .csv) contains the correlation analysis between RQ1 and RQ2 results as discussed in the "Threats to construct validity" (Section 6).</li> <li><strong>rq3-inter-rater</strong> (.tex and .csv) contains the results of the inter-rater agreements analysis discussed in Section 3.6.</li> </ul>
Figure 23. Strict consensus cladograms constructed using a in Identification of fossil worm tubes from Phanerozoic hydrothermal vents and cold seeps
Figure 23. Strict consensus cladograms constructed using a total of 64 modern and fossil annelid taxa and 48 mostly morphological tube characters. Analyses were performed using implied character weighting, with the concavity constant set as default (k = 3; A), and also set to downweight homoplastic characters less (k = 4; B). Numbers on nodes represent groups present/contradicted support values. Modern taxa are coloured according to taxonomic groups; fossil taxa are in grey. A, consensus of 271 most parsimonious trees (best score = 15.387, consistency index = 0.195, retention index = 0.264); B, consensus of 60 most parsimonious trees (best score = 13.568, consistency index = 0.232, retention index = 0.569). Symbols/colours indicate taxonomic affinities.
Figure 1 in Construction of a phylogenetic matrix: Scripts and guidelines for phylogenomics
Figure 1. Flowchart of constructing a phylogenetic matrix for phylogenomics. The custom scripts used in each step are marked as italic. Dashed boxes indicate that these strategies of each step choose only one or more suitable strategies.
Fig. 1 in Changes in soil moisture and riparian forest structure after a dam construction
Fig. 1. Satellite image of a riparian forest on southern Brazil. Study area image with square showing plots locations. A = Spillway and the beginning of Reduced Outflow Stretch, A' = end of Reduced Outflow Stretch, B = hydroeletric dam, B' = end of hydroelectric dam, C = artificial lake created by dam, D = river patch returns to normal flow. The square ilustrates the study area.
Fig. 3 in Changes in soil moisture and riparian forest structure after a dam construction
Fig. 3. Major changes that drives the community changes. Before river diversion, the sectors near the river had greater basal areas because they had many thick trees while distant sectors had thin trees (the density was statistically similar). After four years of river diversion, there were many trunks of still alive trees and dead trees in the sector closer to the river. Even with high growth, the basal area in this sector was severely reduced and became similar to the distant sector (which already has small basal area).
Fig. 2 in Changes in soil moisture and riparian forest structure after a dam construction
Fig. 2. Soil moisture changes that occurred due to construction of the dams. A and C represent soil moisture in dry forests before damming, and B and D represent soil moisture after damming construction. The continuous line represents soil surface; vertical black bars represent soil sampling sites; blue bars represent soil moisture and their thickness illustrates soil moisture; and thicker bars represent more moisture. After dam influence, soil moisture increased mainly in the dry season and mainly near the lakeshore.
1964 flamenca negra guitar from Faustino Conde - documentation and reverse engineered construction plans
<p>This data set documents a specific guitar made by Faustino Conde in Madrid in 1964.<br>(Comment: Felipe Conde identified the handwriting of Mariano Conde in the "Conde" signature on the label inside, the signature of his father. This does not necessarily mean that the guitar was built by Mariano, since experts believe, that it was the brother Faustino who built these kind of special guitars, as it has also been the case for the flamenca negra guitars).</p> <p>The guitar is a so-called flamenca negra, a flamenco guitar with rosewood used for rib and back plate.<br>(Errata: the guitar was sold as a flamenca negra, and experts in Granada also believed it is a flamenca negra. However, Felipe Conde now inspected the guitar and states that the instrument has been built as a classical guitar with only some changes in the setup that might lead to other conclusions. It is not clear whether this setup is original from Conde or whether the guitar has been changed in an aftermath.)</p> <p>The guitar has the exceptional character of showing four main resonances between the Helmholtz resonance (A0) and the first air mode (A1), while other guitars usually have one or two resonances in the same range.<br>This is observable across a wider range of old and contemporary guitars, documented in an archive.<br>Mores, R. 2021a. ‘Archive for the acoustical documentation of classical Spanish guitars, flamenco guitars and romantic guitars from private and public collections – bridge mobility’. <em>Zenodo</em>,<br><a href="https://doi.org/10.5281/zenodo.4604577" target="_blank" rel="noopener">doi: 10.5281/zenodo.4604577</a>.<br>This observation triggered a research project to understand this.</p> <p>The resulting analytical model reveals the delicate tuning of related parameters in the construction, published in the Journal of the Acoustical Society of America, JASA. <br>Mores, R. 2021. ‘Sound tuning in asymmetrically braced guitars’. <em>J. Acoust. Soc. Am.</em> 149(2), 1041–1057.<br><a href="https://doi.org/10.1121/10.0003378" target="_blank" rel="noopener">https://doi.org/10.1121/10.0003378</a>.<br>This model explains how to design multiple resonances into a guitar so that the fundamental tone is suported for every semitone played. The model matches with findings not only of this Conde guitar but explains tuning issues in general. There should be a translation into Spanish in due time.</p> <p>A brief talk (English) explains the main issues and demonstrates the congruence between the analysis and an mechanical model, build for demonstration purposes.<br>Mores, R. 2020a. ‘Tuning signature modes in guitars - a lesson by Faustino Conde’. <em>Zenodo</em>, <br><a href="https://doi.org/10.5281/zenodo.4624826" target="_blank" rel="noopener">doi: 10.5281/zenodo.4624826</a>.</p> <p>Supporting material (coded in MATLAB) allows researchers and guitar makers to explore the issue.<br>Mores, R. 2020. ‘Tuning asymmetrically braced guitars - analytical models and MATLAB code’, <em>Zenodo</em>, <br><a href="https://doi.org/10.5281/zenodo.4010596" target="_blank" rel="noopener">doi: 10.5281/zenodo.4010596</a>.</p> <p>This publication documents the construction plans of the Faustino Conde guitar.<br>Guitar makers asked for these plans to understand the principles of parameter tuning based on the construction. The documentation comprises:<br>1. construction plans (high resolution)<br>2. photos of the total instrument<br>3. photos of details inside </p> <p>***</p> <p>Comments on the guitar for those who consider to build this guitar model.<br>1. The guitar is original. Also the top and the mechanics. Experts in Granada, while inspecting the varnish, believed that the top plate might have been modified. However, Felipe Conde (*1959) states that the guitar is original. This also includes the fretboard. The fretboard height is declining towards the body and this caused questions by Granadian guitar makers. However, in the workshop of the Conde family several inspected guitars from the 50s through the 70s revealed a likewise decline of the height.<br>2. Felipe Conde inspected the guitar. He believes that the guitar is an experimental guitar and that it has been built as a customized guitar for a highly professional musician. But there are no records on this.<br>3. The setup of the guitar has been modified. The height of the bridge is lowered where the bone sits. This can be inspected in the construction plan. It is not known whether this was done by Conde himself or by someone else in an aftermath. Experts in Granada but also members of the Conde family state that the present setup is perfect.</p>
Generic Constructions Datasets
<p>Add Basic Results from ProSenseAIR Algorithm (https://github.com/JanChristianRedlich/ProSenseAIR) on NOW Corpus inclusive Visualizations.</p>
Selected properties and microstructure of concrete with tire rubber granulate as recycled material in construction industry
<p><span>The paper explores the use of recycled materials in the construction industry to promote sustainable development. There is a growing demand for recycling and innovative materials in engineering. The study specifically investigates the potential of tire rubber recyclate as a recycled raw material, comparing two different mixtures in an experimental program. These mixtures highlight the importance of utilizing local resources, aligning with the principles of the circular economy. The experimental program focuses on evaluation of mechanical properties in addition to specialized tests. Findings indicate that higher proportions of rubber granulate not only impact mechanical properties but also significantly affect durability when exposed to environmental factors. </span></p>
Resources of IncRML: Incremental Knowledge Graph Construction from Heterogeneous Data Sources
<h2>IncRML resources</h2> <p>This Zenodo dataset contains all the resources of the paper 'IncRML: Incremental Knowledge Graph Construction from Heterogeneous Data Sources' submitted to the Semantic Web Journal's Special Issue on Knowledge Graph Construction. This resource aims to make the paper experiments fully reproducible through our <a href="https://github.com/kg-construct/exectool" target="_blank" rel="noopener">experiment tool</a> written in Python which was already used before in the <a href="https://doi.org/10.5281/zenodo.7837289" target="_blank" rel="noopener">Knowledge Graph Construction Challenge by the ESWC 2023 Workshop on Knowledge Graph Construction</a>. The exact Java JAR file of the RMLMapper (rmlmapper.jar) is also provided in this dataset which was used to execute the experiments. This JAR file was executed with Java OpenJDK 11.0.20.1 on Ubuntu 22.04.1 LTS (Linux 5.15.0-53-generic). Each experiment was executed 5 times and the median values are reported together with the standard deviation of the measurements.</p> <h2>Datasets</h2> <p>We provide both dataset dumps of the GTFS-Madrid-Benchmark and of real-life use cases from Open Data in Belgium.<br>GTFS-Madrid-Benchmark dumps are used to analyze the impact on execution time and resources, while the real-life use cases aim to verify the approach on different types of datasets since the GTFS-Madrid-Benchmark is a single type of dataset which does not advertise changes at all.</p> <h3>Benchmarks</h3> <ul> <li>GTFS-Madrid-Benchmark: change types with fixed data size and amount of changes: additions-only, modifications-only, deletions-only (11 versions)</li> <li>GTFS-Madrid-Benchmark: amount of changes with fixed data size: 0%, 25%, 50%, 75%, and 100% changes (11 versions)</li> <li>GTFS-Madrid-Benchmark: data size with fixed amount of changes: scales 1, 10, 100 (11 versions)</li> </ul> <h3>Real-world datasets</h3> <ul> <li>Traffic control center Vlaams Verkeerscentrum (Belgium): traffic board messages data (1 day, 28760 versions)</li> <li>Meteorological institute KMI (Belgium): weather sensor data (1 day, 144 versions)</li> <li>Public transport agency NMBS (Belgium): train schedule data (1 week, 7 versions)</li> <li>Public transport agency De Lijn (Belgium): busses schedule data (1 week, 7 versions)</li> <li>Bike-sharing company BlueBike (Belgium): bike-sharing availability data (1 day, 1440 versions)</li> <li>Bike-sharing company JCDecaux (EU): bike-sharing availability data (1 day, 1440 versions)</li> <li>OpenStreetMap (World): geographical map data (1 day, 1440 versions)</li> </ul> <h3>Ingestion</h3> <p>Real-world datasets LDES output was converted into SPARQL UPDATE queries and executed against Virtuoso to have an estimate for non-LDES clients how incremental generation impacted ingestion into triplestores.</p> <h2>Remarks</h2> <ol> <li>The first version of each dataset is always used as a baseline. All next versions are applied as an update on the existing version. The reported results are only focusing on the updates since these are the actual incremental generation.</li> <li>GTFS-Change-50_percent-{ALL, CHANGE}.tar.xz datasets are not uploaded as GTFS-Madrid-Benchmark scale 100 because both share the same parameters (50% changes, scale 100). Please use GTFS-Scale-100-{ALL, CHANGE}.tar.xz for GTFS-Change-50_percent-{ALL, CHANGE}.tar.xz</li> <li>All datasets are compressed with XZ and provided as a TAR archive, be aware that you need sufficient space to decompress these archives! 2 TB of free space is advised to decompress all benchmarks and use cases. The expected output is provided as a ZIP file in each TAR archive, decompressing these requires even more space (4 TB).</li> </ol> <h2>Reproducing</h2> <p>By using our <a href="https://github.com/kg-construct/exectool" target="_blank" rel="noopener">experiment tool</a>, you can easily reproduce the experiments as followed:</p> <ol> <li>Download one of the TAR.XZ archives and unpack them.</li> <li>Clone the GitHub repository of our experiment tool and install the Python dependencies with '<em>pip install -r requirements.txt'.</em></li> <li>Download the rmlmapper.jar JAR file from this Zenodo dataset and place it inside the experiment tool root folder.</li> <li>Execute the tool by running: '<em>./exectool --root=/path/to/the/root/of/the/tarxz/archive --runs=5 run</em>'. The argument '<em>--runs=5</em>' is used to perform the experiment 5 times.</li> <li>Once executed, you can generate the statistics by running: '<em>./exectool --root=/path/to/the/root/of/the/tarxz/archive stats</em>'.</li> </ol> <h2>Testcases</h2> <p>Testcases to verify the integration of RML and LDES with IncRML, see <a href="https://doi.org/10.5281/zenodo.10171394">https://doi.org/10.5281/zenodo.10171394</a></p>
Person Detection on Construction Sites
<p>All information can be found here:</p> <p><a href="https://github.com/aidresden/person_detection_construction_sites" target="_blank" rel="noopener">https://github.com/aidresden/person_detection_construction_sites</a></p> <p>The development of this data set was funded by the Federal Ministry of Labour and Social Affairs (BMAS) and the Federal Institute for Occupational Safety and Health (BAuA) under the administrative agreement ‘Artificial Intelligence in a Safe and Healthy Working Environment’.</p>
Data supporting "A comprehensive analysis of air-sea CO2 flux uncertainties constructed from surface ocean data products"
<p>Changelog</p> <p>v2: Fixes an identified issue in FluxEngine v4.0.7 that affects the calculation of fCO2atm. Fluxes have been recalculated using FluxEngine v4.0.9.1, and the analysis regenerated. The intergrated air-sea CO2 flux (or ocean sink) has reduced by ~0.2-0.3Pg C yr-1 but uncertainties are unchanged. </p> <p>v1: Initial dataset released along with the supporting manuscript</p> <p> </p> <p>Data included in this repository supports the manuscript "A comprehensive analysis of air-sea CO<sub>2</sub> flux uncertainties constructed from surface ocean data products".</p> <p>Two files are present:</p> <ol> <li>A Python config file used to run the software developed for the analysis (Ford et al., 2024)</li> <li>A ZIP file containing the input, neural network, and output files for the analysis.</li> </ol> <p>Within the ZIP file, multiple folders are present:</p> <ol> <li>Decorrelation contains .csv files that contain the annual estimates of the decorrelation lengths for the parameters requiring these (SST, sea ice, wind, fCO<sub>2</sub> and fCO<sub>2</sub> network).</li> <li>Flux contains the individual FluxEngine output files that provide all the flux calculations, and auxillary data to the flux calculations.</li> <li>Fluxengine_input contains the input files to FluxEngine, which specifies the fCO<sub>2 (sw), </sub>xCO<sub>2 (atm)</sub> and the temperature, salinities for the skin and subskin layers.</li> <li>Inputs contains all the monthly 1 degree input data used. Many of the data used are not native monthly 1 deg, and so these are generated from the higher resolution data. These are all combined into the neural_network_input.nc file, so a single file can be distributed with all the inputs used.</li> <li>Networks contains the TensorFlow neural network (FNN) files, where each province has 10 folders (one for each ensemble).</li> <li>Plots contains output plots for debugging and final plots of uncertainties</li> <li>Scalars contains the scalars used to normalise the data before input into the neural network. These are saved as Python pickle files, as they are needed if the neural network is used on other data.</li> <li>Unc_lut contains the look up tables to generate the parameter uncertainty as described in the manuscript. These are Python pickle files.</li> <li>Validation contains a csv file with the independent test RMSD, along with Python Pickle files of the validation data.</li> </ol> <p>In the main folder, three files are present:</p> <ol> <li>Annual_flux.csv contains the annual air-sea CO<sub>2</sub> flux (or ocean sink estimate) estimated from the fCO<sub>2 (sw)</sub> fields. This also contains the annual integrated uncertainties for each component in the uncertainty flow chart in the manuscript.</li> <li>Output.nc contrains the gridded global fields of the fCO<sub>2 (sw)</sub>, the air-sea CO<sub>2</sub> flux, and the uncertainties for all the individual components. Metadata within the file should provide all the information required.</li> <li>Training.tsv contains the training/validation data alongside the input parameters for neural network training</li> </ol> <p> </p> <p>Please contact Daniel J. Ford (<a href="mailto:d.ford@exeter.ac.uk">d.ford@exeter.ac.uk</a>) if you have any questions.</p> <p><strong>Acknowledgements</strong></p> <p>This work was funded by the Convex Seascape Survey (https://convexseascapesurvey.com/) and the European Union under grant agreement no. 101083922 (OceanICU; https://ocean-icu.eu/) and UK Research and Innovation (UKRI) under the UK government’s Horizon Europe funding guarantee [grant number 10054454, 10063673, 10064020, 10059241, 10079684, 10059012, 10048179]. The views, opinions and practices used to produce this dataset/software are however those of the author(s) only and do not necessarily reflect those of the European Union or European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.</p> <p>The Surface Ocean CO₂ Atlas (SOCAT) is an international effort, endorsed by the International Ocean Carbon Coordination Project (IOCCP), the Surface Ocean Lower Atmosphere Study (SOLAS) and the Integrated Marine Biosphere Research (IMBeR) program, to deliver a uniformly quality-controlled surface ocean CO₂ database. The many researchers and funding agencies responsible for the collection of data and quality control are thanked for their contributions to SOCAT.</p> <p> </p> <p><strong>References</strong></p> <p>Ford, D. J., Blannin, J., Watts, J., Watson, A. J., Landschutzer, P., Jersild, A., & Shutler, J. D. (2024, June 30). OceanICU Neural Network Framework with per pixel uncertainty propagation (v1.1) (Version v1.1). Zenodo. https://doi.org/10.5281/ZENODO.12597803</p>
Mapping 'the constructive turn' in comment sections of news websites
<p>The research project critically examines the guidelines of the comments sections of the twenty largest online news outlets over the last ten years. Rather than focusing on the familiar negative comments of news consumers and their narratives, we analyze and compare the news outlets’ guidelines and how they have led in what we call ‘a constructive turn’. We propose our own theoretical framework to analyze what is encouraged and what is discouraged in news outlets’ guidelines. Results show an increasing focus on constructiveness in the guidelines of the comment sections and a shift to more positivity, rather than on deleting and filtering negative or toxic comments. Although platforms differ in their views on the role of commenting and the definition of constructiveness, the turn towards the constructive design of the commenting platform is shared among them.</p> <p>This dataset contains the commentary guidelines in the top 20 English-language online news websites of December 2020 based on research conducted by Similar Web (Source: Similar Web for Gazette). For each news publication, the current commentary guidelines were scrapped from the internet, alongside earlier versions of their guidelines. In total, three moments were used to map the guidelines: 2021, 2015 and 2010. The content was analysed through coding using Nvivo software. We applied a bottom-up approach - by creating simple codes and eventually grouping them together. Each set of guidelines was coded on what behaviour was <em>encouraged</em> and what was <em>discouraged</em> by the news outlet, and what kind of <em>discussion environment </em>the news outlet expects from their commenters in general (e.g. entertaining, healthy, inclusive etc.). </p> <p>This dataset contains coded content for the project. Following logic was used in uploading the documents:</p> <p>1 - Nvivo project file - can be opened using Nvivo for Mac - contains all information (files, codes, etc.)</p> <p>We also upload more user-friendly data (the following documents are uploaded in MS Word format):</p> <p>2 - Codebook (provides the logical structure of coding applied + number of codes for each category)<br> 3 - Code excerpts for discouraged elements found in the content<br> 4 - Code excerpts for encouraged elements found in the content<br> 5 - Code excerpts for discussion environment elements found in the content</p> <p>Disclaimer: The user-generated content guidelines of news media companies are their own intellectual property and we do not own any rights to it. </p> <p><br> </p>
Construction shelter assembly at Coventry University
<p>Timelapse video of the assembly process for the construction shelter designed in ECOBULK and manufactured with circular-economy materials.</p>
TReNCo: Topologically associating domain (TAD) aware regulatory network construction (extended data)
<p>The enclosed files contain all of the extended data from: TReNCo: Topologically associating domain (TAD) aware regulatory network construction</p>
MCPNet : A parallel maximum capacity-based genome-scale gene network construction framework
<p>This deposit contains the gene expression profile datasets used for the paper titled "MCPNet : A parallel maximum capacity-based genome-scale gene network construction framework". </p> <p>There are three sets of data:</p> <ul> <li>Simulated yeast data from NetBenchmark, with random noise injected, as well as the ground truth network matrix. In "SimulatedYeast.zip".</li> <li>Real Yeast dataset and the ground truth network as an adjacency list file. in "yeast_data.exp" and "yeast_gs1_list_filtered.tsv". Data acquired from "Castro DM, de Veaux NR, Miraldi ER, Bonneau R (2019) Multi-study inference of regulatory networks for more accurate models of gene regulation. PLoS Comput Biol 15(1): e1006591. https://doi.org/10.1371/journal.pcbi.1006591", <a href="https://github.com/simonsfoundation/multitask_inferelator/tree/AMuSR">https://github.com/simonsfoundation/multitask_inferelator/tree/AMuSR</a>.</li> <li>Real Arabidopsis athaliana datasets for 5 tissues and 1 environmental challenge. <ul> <li>athaliana_gs_probes.tsv : ground truth as an adjacency list</li> <li>microarray gene expression profiles for "development", "leaf", "seed", "flower", "seedling1week", "hormone-aba-iaa-ga-br".</li> <li>Aathaliana.Datasets-CEL-File-URLs.xlsx: list of SRA accession numbers for the A. athaliana datasets</li> </ul> </li> </ul>
Data from: Improving quartet graph construction for scalable and accurate species tree estimation from gene trees
<p>Summary methods are one of the dominant approaches for estimating species trees from genome-scale data. However, they can fail to produce accurate species trees when the input gene trees are highly discordant due to gene tree estimation error as well as biological processes, like incomplete lineage sorting. Here, we introduce a new summary method TREE-QMC that offers improved accuracy and scalability under these challenging scenarios. TREE-QMC builds upon the algorithmic framework of QMC (Snir and Rao 2010) and its weighted version wQMC (Avni et al. 2014). Their approach takes weighted quartets (four-leaf trees) as input and builds a species tree in a divide-and-conquer fashion, at each step constructing a graph and seeking its max cut. We improve upon this methodology in two ways. First, we address scalability by providing an algorithm to construct the graph directly from the input gene trees. By skipping the quartet weighting step, TREE-QMC has a time complexity of O(n^3 k) with some assumptions on subproblem sizes, where n is the number of species and k is the number of gene trees. Second, we address accuracy by normalizing the quartet weights to account for "artificial taxa," which are introduced during the divide phase so that solutions on subproblems can be combined during the conquer phase. Together, these contributions enable TREE-QMC to outperform the leading methods (ASTRAL-III, FASTRAL, wQFM) in an extensive simulation study. We also present the application of these methods to an avian phylogenomics data set.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.