Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
166
datasets available to search
ShareScore release 0.7.1
Dataset results
166 results for “Open-source”
Dataset related to the manuscript: "An open-source integrated framework for the automation of citation collection and screening in systematic reviews"
<p>Dataset related to the manuscript: “An open-source integrated framework for the automation of citation collection and screening in systematic reviews”, to be used together with the code stored at https://github.com/AD-Papers-Material/BART_SystReviewClassifier to reproduce the results.</p> <p>There are three datasets:<br> - The Record data collected from the online scientific databases;<br> - The session journal which describes the search session, i.e., how many records were collected and from which source, for each query/session pairs.<br> - The session data which is the outcome of the classification and review tasks;</p>
OpenFOAM cases of the paper "Development and validation of an open-source CFD model for the efficiency assessment of data centers"
<p>This dataset contains the<em> underling data</em> for the paper "Development and validation of an open-source CFD model for the efficiency assessment of data centers”, submitted for the consideration and open review in Open Research Europe (ORE).</p> <p><strong>Validation1.tar.xz:</strong> OpenFOAM files and scripts for the simulation of flow and thermal structures in an enclosed environment (Wang and Chen, 2009).</p> <p><em>Wang, Miao; Chen, Qingyan (2009). Assessment of Various Turbulence Models for Transitional Flows in an Enclosed Environment (RP-1271). HVAC&R Research, 15(6), 1099–1119. doi:10.1080/10789669.2009.10390881</em></p> <p><strong>Validation2-kOmegaSSTModel.tar.xz:</strong> OpenFOAM files and scripts for the simulation of forced convection in a room (Zhang et al. 2007) using k-omega SST turbulence model. </p> <p><em>Zhao Zhang, Wei Zhang, Zhiqiang John Zhai & Qingyan Yan Chen (2007) Evaluation of Various Turbulence Models in Predicting Airflow and Turbulence in Enclosed Environments by CFD: Part 2—Comparison with Experimental Data from Literature, HVAC&R Research, 13:6, 871-886, DOI: 10.1080/10789669.2007.10391460</em></p> <p><strong>Validation2-RNGkEpsilonModel.tar.xz:</strong> OpenFOAM files and scripts for the simulation of forced convection in a room (Zhang et al. 2007) using RNG k-epsilon turbulence model. </p> <p><em>Zhao Zhang, Wei Zhang, Zhiqiang John Zhai & Qingyan Yan Chen (2007) Evaluation of Various Turbulence Models in Predicting Airflow and Turbulence in Enclosed Environments by CFD: Part 2—Comparison with Experimental Data from Literature, HVAC&R Research, 13:6, 871-886, DOI: 10.1080/10789669.2007.10391460</em></p> <p><strong>Validation3.tar.xz:</strong> OpenFOAM files and scripts for the simulation of strong natural convection in a model fire room (Murakami et al. 1995).</p> <p><em>Murakami, S., S. Kato, and R. Yoshie. 1995. Measurement of turbulence statistics in a model fire room by LDV. ASHRAE Transactions 101(2):287–301.</em></p> <p><strong>Validation4.tar.xz:</strong> OpenFOAM files and scripts for the simulation of thermal distribution in an open-aisle data center (Abdelmaksoud et al. 2013).</p> <p><em>W.A. Abdelmaksoud, T.Q. Dang, H. Ezzat Khalifa, R.R. Schmidt Improved computational fluid dynamics model for open-aisle air-cooled data center simulations J. Electron. Packag., 135 (2013), pp. 030901-30913</em></p> <p><strong>Results_Validation1.tar.xz:</strong> Simulation results of the Validation case 1.</p> <p><strong>Results_Validation2.tar.xz:</strong> Simulation results of the Validation case 2.</p> <p><strong>Results_Validation3.tar.xz:</strong> Simulation results of the Validation case 3.</p> <p><strong>Results_Validation4.tar.xz:</strong> Simulation results of the Validation case 4.</p> <p><strong>layout.csv:</strong> Input file for the Validation case 4.</p>
AN OPEN-SOURCE, THREE-DIMENSIONAL GROWTH MODEL OF THE MANDIBLE
<p>This repository contains all geometrical data and metadata belonging to the paper AN OPEN-SOURCE, THREE-DIMENSIONAL GROWTH MODEL OF THE MANDIBLE by the MAGIC Amsterdam research consortium. The following contents are uploaded:</p><p><strong>shapeVectors_original.csv</strong> | shape vectors of the original data<br><strong>shapeVectors_rescaled.csv</strong> | shape vectors of the rescaled data<br>678 x 62589 matrices where the rows are samples and the columns are shape vectors. The shape vectors are formatted<i> [x1, x2, x3, ..., y1, y2, y3, ..., z1, z2, z3, ...].</i></p><p><strong>PCA_coeff_original.csv</strong> | principal component coefficients of the original data<br><strong>PCA_coeff_rescaled.csv</strong> | principal component coefficients of the rescaled data<br>62589 x 677 matrices where each row of these matrices is a variable (x-, y-, or z-coordinate of a vertex) and each column is a principal component.</p><p><strong>PCA_score_original.csv</strong> | principal component scores of the original data<br><strong>PCA_score_rescaled.csv</strong> | principal component scores of the rescaled data<br>678 x 677 matrices where rows correspond to samples and columns correspond to principal components.</p><p><strong>PCA_latent_original.csv</strong> | principal component variances of the original data<br><strong>PCA_latent_rescaled.csv</strong> | principal component variances of the rescaled data<br>677 x 1 vectors where each element is an eigenvalue of a principal component.</p><p><strong>PCA_mu_original.csv</strong> | mean of the original data<br><strong>PCA_mu_rescaled.csv</strong> | mean of the rescaled data<br>1 x 62589 vectors that represent the average shape vector. All (centered) data can be reconstructed as follows: <i>shapeVectors = PCA_score * PCA_coeff' + PCA_mu.</i></p><p><strong>PCA_standardDeviations_original.csv</strong> | standard deviations of each sample for each principal component of the original data.<br><strong>PCA_standardDeviations_rescaled.csv</strong> | standard deviations of each sample for each principal component of the rescaled data.<br>677 x 678 matrices where the rows are principal components and the columns are samples. The standard deviations were calculated as follows: <i>PCA_standardDeviations = PCA_score' ./ sqrt(PCA_latent).</i></p><p><strong>metadata.csv</strong> | This matrix contains the age in years (first column) and biological sex (second column, 1 = male and 2 = female) for all samples (rows).</p><p><strong>connectivityList.csv</strong> | This matrix defines the mesh of the 3D model of the mandible. The vector in each row represents which vertices define a triangle. Indexing starts at 0, so for use in e.g. Matlab, add 1 to all elements.</p>
Open-source traffic and CO2 emission dataset for commercial aviation
<p>This record is a global open-source passenger air traffic dataset primarily dedicated to the research community. <br>It gives a seating capacity available on each origin-destination route for a given year, 2019, and the associated aircraft and airline when this information is available. </p> <p>Context on the original work is given in the related articles (<a href="https://doi.org/10.59490/joas.2024.7365">https://doi.org/10.59490/joas.2024.7365,</a> <a href="https://doi.org/10.59490/joas.2023.7201">https://doi.org/10.59490/joas.2023.7201)</a> and on the associated GitHub page (<a href="https://github.com/AeroMAPS/AeroSCOPE/">https://github.com/AeroMAPS/AeroSCOPE/</a>).<br>A simple data exploration interface will be available at <a href="www.aeromaps.eu/aeroscope">www.aeromaps.eu/aeroscope.</a><br>The dataset was created by aggregating various available open-source databases with limited geographical coverage. It was then completed using a route database created by parsing Wikipedia and Wikidata, on which the traffic volume was estimated using a machine learning algorithm (XGBoost) trained using traffic and socio-economical data.<br> </p> <h4><br><strong>1- DISCLAIMER</strong></h4> <p><br>The dataset was gathered to allow highly aggregated analyses of the air traffic, at the continental or country levels. At the route level, the accuracy is limited as mentioned in the associated article and improper usage could lead to erroneous analyses. </p> <p>Although all sources used are open to everyone, the Eurocontrol database is only freely available to academic researchers. It is used in this dataset in a very aggregated way and under several levels of abstraction. As a result, it is not distributed in its original format as specified in the contract of use.</p> <p>As a general rule, we decline any responsibility for any use that is contrary to the terms and conditions of the various sources that are used. In case of commercial use of the database, please contact us in advance.</p> <h4><br><strong>2- DESCRIPTION</strong></h4> <p>Each data entry represents an (Origin-Destination-Operator-Aircraft type) tuple.</p> <p><em>Please </em>refer<em> to </em>the<em> support article for more details (see above).</em></p> <p>The dataset contains the following columns:</p> <ul> <li>"First column" : index</li> <li><strong>airline_iata : </strong>IATA code of the operator in nominal cases. An ICAO -> IATA code conversion was performed for some sources, and the ICAO code was kept if no match was found.</li> <li><strong>acft_icao : </strong>ICAO code of the aircraft type</li> <li><strong>acft_class : </strong>Aircraft class identifier, own classification. <ul> <li>WB: Wide Body</li> <li>NB: Narrow Body</li> <li>RJ: Regional Jet</li> <li>PJ: Private Jet</li> <li>TP: Turbo Propeller</li> <li>PP: Piston Propeller</li> <li>HE: Helicopter</li> <li>OTHER</li> </ul> </li> <li><strong>seymour_proxy: </strong>Aircraft code for Seymour Surrogate (https://doi.org/10.1016/j.trd.2020.102528), own classification to derive proxy aircraft when nominal aircraft type unavailable in the aircraft performance model.</li> <li><strong>source: </strong>Original data source for the record, before compilation and enrichment. <ul> <li>ANAC: Brasilian Civil Aviation Authorities</li> <li>AUS Stats: Australian Civil Aviation Authorities</li> <li>BTS: US Bureau of Transportation Statistics T100</li> <li>Estimation: Own model, estimation on Wikipedia-parsed route database</li> <li>Eurocontrol: Aggregation and enrichment of R&D database</li> <li>OpenSky</li> <li>World Bank</li> </ul> </li> <li><strong>seats: </strong>Number of seats available for the data entry, AFTER airport residual scaling</li> <li><strong>n_flights: </strong>Number of flights of the data entry, when available</li> <li><strong>iata_departure</strong>, <strong>iata_arrival : </strong>IATA code of the origin and destination airports. Some BTS inhouse identifiers could remain but it is marginal.</li> <li><strong>departure_lon</strong><em>, </em><strong>departure_lat</strong><em>, </em><strong>arrival_lon</strong><em>, </em><strong>arrival_lat : </strong>Origin and destination coordinates, could be NaN if the IATA identifier is erroneous</li> <li><strong>departure_country, arrival_country</strong>: Origin and destination country ISO2 code. <strong>WARNING: </strong>disable NA (Namibia) as default NaN at import</li> <li><strong>departure_continent, arrival_continent: </strong>Origin and destination continent code. <strong>WARNING: </strong>disable NA (North America) as default NaN at import</li> <li><strong>seats_no_est_scaling: </strong>Number of seats available for the data entry, BEFORE airport residual scaling</li> <li><strong>distance_km: </strong>Flight distance (km)</li> <li><strong>ask: </strong>Available Seat Kilometres</li> <li><strong>rpk: </strong>Revenue Passenger Kilometres (simple calculation from ASK using IATA average load factor)</li> <li><strong>fuel_burn_seymour: </strong>Fuel burn <em>per flight</em> (kg) when seymour proxy available</li> <li><strong>fuel_burn: </strong>Total fuel burn of the data entry (kg)</li> <li><strong>co2: </strong>Total CO2 emissions of the data entry (kg)</li> <li><strong>domestic: </strong>Domestic/international boolean (Domestic=1, International=0)</li> </ul> <p> </p> <h4><strong>3- Citation</strong></h4> <p>Please cite the support paper instead of the dataset itself. </p> <blockquote> <p>Salgas, A., Sun, J., Delbecq, S., Planès, T., & Lafforgue, G. (2024). Compilation and Applications of an Open-Source Dataset on Global Air Traffic Flows and Carbon Emissions. <em>Journal of Open Aviation Science</em>. <a href="https://doi.org/10.59490/joas.2024.7365">https://doi.org/10.59490/joas.2023.7201</a></p> </blockquote>
MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes
<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here: <a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). </p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2 Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name: </strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe: </strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' values over 50%: FLAG_RP63.</li> <li><strong>flag_sum: </strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value). </p> <p> </p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis. </p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0 functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes. <br><br></p>
NMRduino: A modular, open-source, low-field magnetic resonance platform
<p>The NMRduino is a compact, cost-effective, sub-MHz NMR spectrometer that utilizes readily available open-source hardware and software components. One of its aims is to simplify the processes of instrument setup and data acquisition control to make experimental NMR spectroscopy accessible to a broader audience. In this introductory paper, the key features and potential applications of NMRduino are described to highlight its versatility both for research and education.</p>
Quality-Assurance Package for the "Automated, Open-Source, Vendor-Independent Quality Assurance Protocol Based on the Pulseq Framework" Manuscript
<h2>Background</h2> <p>Neuroimaging research requires consistent image quality and temporal signal stability, especially for functional magnetic resonance imaging (MRI) studies that rely on detecting subtle blood-oxygen-level-dependent (BOLD) signal changes. Regular MR system performance monitoring is essential, especially for longitudinal and multi-site studies. This study aims to establish a robust quality assurance (QA) protocol to promote data comparability across scanner models, vendors, and sites, as well as over a prolonged period.</p> <p>The manuscript titled "<em>Automated, Open-Source, Vendor-Independent Quality Assurance Protocol Based on the Pulseq Framework</em>" was submitted to the Special Issue <a href="https://link.springer.com/journal/10334/updates/26638300">Reproducibility and Quality Assurance</a> of the Magnetic Resonance Materials in Physics, Biology and Medicine (MAGMA) journal.</p> <p>This QA package proposed by the manuscript hosts materials for</p> <ul> <li>all reconstructed images,</li> <li>instruction for data acquisition,</li> <li>instruction for image reconstruction,</li> <li>instruction for post-processing,</li> <li>example raw data and DICOM images, and</li> <li>images and scripts for T1/T2 fitting.</li> </ul> <p>The detailed information is listed below.</p> <h2>All reconstructed images</h2> <p>This directory contains all reconstructed images from the fBIRN phantom on three Siemens 3T scanners (Trio, Prisma.Fit, and Cima.X) and one GE (UHP) 3T scanner. It contains four sub-folders for each scanner. And each sub-folder contains (some of) the following sub-folders:</p> <ul> <li><code>product_epi_ice</code>: ICE-reconstructed product EPI images.</li> <li><code>product_epi_gt</code>: Gadgetron-reconstructed product EPI images.</li> <li><code>pulseq_epi_ice</code>: ICE-reconstructed Pulseq EPI images.</li> <li><code>pulseq_epi_gt</code>: Gadgetron-reconstructed Pulseq EPI images.</li> <li><code>product_se_ice</code>: ICE-reconstructed product spin-echo (SE) images.</li> <li><code>product_se_gt</code>: Gadgetron-reconstructed product SE images.</li> <li><code>pulseq_se_ice</code>: ICE-reconstructed Pulseq SE images.</li> <li><code>pulseq_se_gt</code>: Gadgetron-reconstructed Pulseq SE images.</li> </ul> <h2>Instruction for data acquisition</h2> <p>This directory includes the following documents:</p> <ul> <li><code>write_QA_Tran_EPIrs.m</code> to generate the <code>QA_epi.seq</code> file for EPI scans.</li> <li><code>write_QA_Tran_T1.m</code>: to generate the <code>QA_T1.seq</code> file for SE scans.</li> <li><code>20241122_QA_protocol_instruction_siemens.docx</code>: standard operating procedure for QA measurements.</li> <li><code>QA_record.xlsx</code>: Excel sheet for the record of QA measurements.</li> </ul> <h2>Instruction for image reconstruction</h2> <h3><em>Documents</em></h3> <ul> <li><code>pulseq2mrd_epi.m</code>: convert GE Pulseq EPI raw data (<code>.mat</code>) to MRD raw data (<code>.h5</code>) using the LABEL information in the <code>QA_epi.seq</code> file.</li> <li><code>pulseq2mrd_se.m</code>: convert GE Pulseq SE raw data (<code>.mat</code>) to MRD raw data (<code>.h5</code>) using the LABEL information in the <code>QA_T1.seq</code> file.</li> <li><code>siemens2mrd_epi.m</code>: convert Siemens Pulseq EPI raw data (<code>.dat</code>) to MRD raw data (<code>.h5</code>) using the information in the <code>.dat</code> raw data.</li> </ul> <ul> <li><code>default.xml</code>: Gadgetron configuration file for SE image reconstruction. This document is already in the Gadgetron container: <code>/opt/conda/envs/gadgetron/share/gadgetron/config/default.xml</code>.</li> <li><code>qc_epi.xml</code>: Gadgetron configuration file for EPI image reconstruction, which is modified from the <code>default epi.xml</code> located in the Gadgetron container: <code>/opt/conda/envs/gadgetron/share/gadgetron/config/</code>.</li> </ul> <ul> <li><code>specialCard_ICE.png</code>: Special card setting for ICE online reconstruction.</li> </ul> <h3><em>Procedures for Gadgetron offline reconstruction</em></h3> <p><strong>Step 1: Gadgetron installation (for more details, visit <a href="https://gadgetron.github.io/tutorial/">here</a>)</strong></p> <ul> <li>Download and install <a href="https://www.docker.com/">Docker</a> software. You may need to install/update the Windows Sub Linux (WSL) system for the Docker installation.</li> <li>Open your terminal (Power shell with administrative privilege in Windows) and navigate to the folder you would like to map to the Gadgetron Docker container.</li> <li>Run: <code>docker run -t --name gt_latest --detach --volume ${pwd}:/opt/data ghcr.io/gadgetron/gadgetron/gadgetron_ubuntu_rt_nocuda:latest</code>. If docker is not recognized, set <code>docker</code> to connect to <code>C:\Program Files\Docker\Docker\resources\bin</code> in the Environment Path in Windows. This will download and then launch the <a href="https://gadgetron.readthedocs.io/en/latest/building.html">latest Gadgetron version</a> in a Docker container. It will also mount your current folder as a data folder inside the container.</li> <li>Run this command: <code>docker exec -ti gt_latest /bin/bash</code>. This will execute your Gadgetron container.</li> </ul> <p><strong>Step 2: Data preparation</strong></p> <ul> <li>Place your SE/EPI <code>.dat</code>/<code>.h5</code> data in the mounted folder.</li> <li>Run the command in Terminal: <code>cd /opt/data</code> to enter the mounted folder.</li> </ul> <p><strong>Step 3: MRD conversion</strong></p> <ul> <li>For Siemens data, you can convert the <code>.dat</code> data to MRD data by using Gadgetron. If Gsdgetron doesn't work (e.g. for XA EPI data), you can then use the Matlab script <code>siemens2mrd_epi.m</code>.</li> <li>The command for Siemens SE data conversion: <code>siemens_to_ismrmrd -f meas_MID*.dat -z 2 -o se_data.h5</code>.</li> <li>The command for Siemens EPI data conversion: <code>siemens_to_ismrmrd -f meas_MID*.dat -z 2 -m IsmrmrdParameterMap_Siemens.xml -x IsmrmrdParameterMap_Siemens_EPI.xsl -o epi_data.h5</code>.</li> <li>For GE data, you can convert the <code>.mat</code> raw data to MRD data by using the Matlab scripts with the corresponding <code>.seq</code> files. For SE conversion: use <code>pulseq2mrd_se.m</code> with <code>QA_T1.seq</code>. For EPI conversion: use <code>pulseq2mrd_epi.m</code> with <code>QA_epi.seq</code>.</li> </ul> <p><strong>Step 4: Gadgetron reconstruction</strong></p> <ul> <li>SE reconstruction: <code>gadgetron_ismrmrd_client -f se_data.h5 -c default.xml -o se_out.h5</code>.</li> <li>EPI reconstruction: first, put <code>qc_epi.xml</code> to the mounted folder and then copy it to the Gadgetron container: <code>cp /opt/data/qc_epi.xml /opt/conda/envs/gadgetron/share/gadgetron/config/</code>. Then, run the reconstruction: <code>gadgetron_ismrmrd_client -f epi_data.h5 -c qc_epi.xml -o epi_out.h5</code>.</li> </ul> <p><strong>Step 5: Load Gadgetron-reconstructed images (<code>.h5</code>)</strong></p> <ul> <li>Load SE <code>.h5</code> images in Matlab:</li> </ul> <blockquote> <p>filename = 'pulseq_se_out.h5' ;</p> <p>info = hdf5info(filename) ;</p> <p>address_data_1 = info.GroupHierarchy.Groups(1).Groups.Datasets(2).Name ;</p> <p>pulseq_se_im = squeeze(double( hdf5read(filename, address_data_1) ) ) ;</p> <p>pulseq_se_im = reshape(pulseq_se_im, [256, 256, 11, 2]) ;</p> </blockquote> <ul> <li>Load EPI <code>.h5</code> images in Matlab:</li> </ul> <blockquote> <p>filename = 'pulseq_epi_out.h5';</p> <p>info = hdf5info(filename) ;</p> <p>address_data_1 = info.GroupHierarchy.Groups(1).Groups.Datasets(2).Name ;</p> <p>pulseq_epi_im = squeeze(double( hdf5read(filename, address_data_1) ) ) ;</p> <p>pulseq_epi_im = reshape(pulseq_epi_im, [64, 64, 27, 200]) ;</p> </blockquote> <h3><em>Procedures for ICE online reconstruction</em></h3> <p>Before executing the Pulseq-based sequences, you can enable ICE online Reconstruction following the procedures below:</p> <ul> <li>Navigate to the Special Card (<code>specialCard_ICE.png</code>), set <code>Data handling</code> to <code>ICE STD</code> for NUMARIS/X (e.g. XA60A and XA61A), and <code>ICE 2D</code> for NUMARIS/4 (e.g. VB, VD, and VE).</li> <li>Select <code>Sum-of-Square</code> for coil combination.</li> <li>Be sure that the maximal pixel intensity does not violate the intensity threshold of <strong>4096</strong>.</li> </ul> <h2>Instruction for post-processing</h2> <p>The example post-processing is based on the reconstructed images from Cima.X over five days.</p> <h3><em>Reconstructed images from Cima.X</em></h3> <p><strong>Note</strong>: All <code>se</code> folders contain a <code>structuralQuality_main.m</code> to call the <code>structuralQuality.m</code> function for structural quality analysis. All <code>epi</code> folders contain a <code>temporalQuality_main.m</code> to call the <code>temporalQuality.m</code> function for temporal quality analysis.</p> <ul> <li><code>product_epi_ice</code>: ICE-reconstructed product EPI images.</li> <li><code>product_epi_gt</code>: Gadgetron-reconstructed product EPI images.</li> <li><code>pulseq_epi_ice</code>: ICE-reconstructed Pulseq EPI images.</li> <li><code>pulseq_epi_gt</code>: Gadgetron-reconstructed Pulseq EPI images.</li> <li><code>product_se_ice</code>: ICE-reconstructed product SE images.</li> <li><code>product_se_gt</code>: Gadgetron-reconstructed product SE images.</li> <li><code>pulseq_se_ice</code>: ICE-reconstructed Pulseq SE images.</li> <li><code>pulseq_se_gt</code>: Gadgetron-reconstructed Pulseq SE images.</li> </ul> <h3><em>QA analysis Matlab package: </em><code><em>QA_functions</em></code></h3> <ul> <li><code>circfit.m</code>: to find the center point and radius of the phantom.</li> <li><code>makeCircleMask.m</code>: to make a circular mask based on the center point and radius.</li> <li><code>structuralQuality.m</code>: to analyze the structural quality of the SE images.</li> <li><code>temporalQuality.m</code>: to analyze the temporal quality of the EPI images.</li> </ul> <h3><em>Post-processing procedures</em></h3> <ul> <li>Step 1: Add the <code>QA_functions</code> folder to your Matlab Path.</li> <li>Step 2: Run the <code>temporalQuality_main.m</code> or <code>structuralQuality_main.m</code> script in each folder to produce the QA results of all reconstructed images inside the folder.</li> <li>Step 3: Run the <code>make_figure_epi.m</code> and <code>make_figure_se.m</code> to produce some of the tables and figures used in the manuscript.</li> </ul> <h2>Example raw data and DICOM images</h2> <p>The data and DICOM images were acquired from Cima.X on the fBIRN phantom on 06.08.2024.</p> <ul> <li>DICOM folder: contains the DICOM images for four EPI scans (the first two scans for warm-up) and two SE scans.</li> <li><code>meas*.dat</code>: Siemens raw data of two EPI scans for temporal quality analysis and two SE scans for structural quality analysis.</li> <li><code>*data.h5</code> files: the ISMRMRD data of the four raw datasets.</li> <li><code>*out.h5</code> files: the images reconstructed by Gadgetron.</li> <li><code>*.nii</code>: the NIFTI-format reconstructed images.</li> <li><code>siemens2mrd_epi.m</code>: to convert the Siemens EPI raw data to ISMRMRD data.</li> <li><code>read_image.m</code>: to convert the Gadgetron-reconstructed h5-format images to NIFTI-format images.</li> </ul> <h2>Images and scripts for T1/T2 fitting</h2> <p>This package includes DICOM images and T1/T2 fitting scripts for the fBIRN phantom. Images for T1 fitting were acquired using a product turbo spin echo sequence with an inversion recovery pulse (repetition time = 4000 ms, echo train length = 4). Images for T2 fitting were obtained using a product SE sequence (repetition time = 3500 ms). Both measurements were conducted on the Siemens Prisma.Fit 3T scanner on 05.06.2024.</p> <ul> <li><code>T1 sub-folder</code>: contains all DICOM images for T1 fitting with inversion recovery times of {50, 150, 300, 450, 600, 750, 900, 1050, 1200, 1350, 1500, 2200, 3000} ms.</li> <li><code>T2 sub-folder</code>: contains all DICOM images for T2 fitting with echo times of {7.5, 15, 30, 45, 60, 75, 90, 130, 200, 250} ms.</li> <li><code>Do_T1fit.m</code>: Matlab script for T1 fitting.</li> <li><code>Do_T2fit.m</code>: Matlab script for T2 fitting.</li> </ul> <p>For more information regarding Pulseq and the workflow for data acquisition and image reconstruction, please visit our GitHub repositories: <a href="https://github.com/pulseq/pulseq">Pulseq Matlab software</a>, <a href="https://github.com/pulseq/tutorials">Pulseq Tutorials</a>, and <a href="https://github.com/pulseq/Pulseq-Rocks-2023-24-ISMRM-Reproducibility-Challenge">Pulseq Rocks for the 2024 ISMRM Reproducibility Team Challenge</a>.</p> <p>If you need any further information or have any questions, please feel free to contact our Pulseq email address: pulseq.mr@uniklinik-freiburg.de.</p>
The potential of low-cost UAVs and open-source photogrammetry software for high-resolution monitoring of alpine glaciers: A case study from the Kanderfirn (Swiss Alps)
<p>This dataset contains high-resolution orthophotos (5 x 5 cm) and digital surface models (25 x 25 cm) of the Kandernfirn Glacier located in the Swiss Alps. Aerial images were aquired with a self-developed fixed-wing Unmanned Aerial Vehicle during ten surveys on five different days in 2017 and 2018. The open-source photogrammetry software OpenDroneMap (version 0.4.1) was used for image processing.</p> <p>The orthophotos and digital surface models were validated through dGNSS point measurements of ground control points. Please refer to the corresponding paper for information on the horizontal and vertical accuracy of the files.</p>
Github commit data for the article "Beyond Zipf's law: Exploring the discrete generalized beta distribution in open-source repositories"
<p><span>This dataframe corresponds to the data used in the Nowak's et al. 2024 article "Beyond Zipf’s law: Exploring the discrete generalized beta distribution in open-source repositories" (see reference below).</span></p> <p><span>It consists of the distirbutions of number of commits per user across a number of GitHub repositories. <br><br>There are three columns:</span></p> <ul> <li><span>repository: the repository name</span></li> <li><span># of commits: the number of commits of a given individual</span></li> <li><span>rank: the user rank in the repository (by decreasing number of commits)<br><br></span></li> </ul> <p><strong><span>Reference:</span></strong></p> <p><span>Nowak, P., Santolini, M., Singh, C., Siudem, G., & Tupikina, L. (2024). Beyond Zipf’s law: Exploring the discrete generalized beta distribution in open-source repositories. <em>Physica A: Statistical Mechanics and Its Applications</em>, <em>649</em>, 129927. <a href="https://doi.org/10.1016/j.physa.2024.129927">https://doi.org/10.1016/j.physa.2024.129927</a></span></p>
Raw and analyzed data for manuscript: "An open-source surface barrier discharge plasma pretreatment for reduced cracking of outdoor wood coatings"
<p><strong>Highlights:</strong></p> <ul> <li>Surface barrier discharges are an affordable and available plasma technology for industrial, laboratory and home-workshop applications.</li> <li>Plasma pretreatments had no impact on the appearance of different protective wood coating for outdoor usage.</li> <li>The weathering performance of outdoor wood coatings improved by plasma, showing less cracks and less biotic factors.</li> </ul>
Towards an open-source landscape for 3D CSEM modelling
<p>Accompanying data to journal article</p> <blockquote> <p>Werthmüller, D., R. Rochlitz, O. Castillo-Reyes, and L. Heagy, 2021, Towards an open-source landscape for 3D CSEM modelling: Geophysical Journal International; ggab238, DOI: <a href="https://doi.org/10.1093/gji/ggab238">10.1093/gji/ggab238</a>.</p> </blockquote> <ul> <li>Official article: <a href="https://doi.org/10.1093/gji/ggab238">https://doi.org/10.1093/gji/ggab238</a></li> <li>GitHub repo: <a href="https://github.com/swung-research/3d-csem-open-source-landscape">https://github.com/swung-research/3d-csem-open-source-landscape</a></li> <li>arXiv.org: <a href="https://arxiv.org/abs/2010.12926">https://arxiv.org/abs/2010.12926</a></li> </ul> <p>The Marlim R3D model can be found at:</p> <ul> <li>Original, fine resistivity model: <a href="https://doi.org/10.5281/zenodo.400233">https://doi.org/10.5281/zenodo.400233</a></li> <li>Upscaled computational model: <a href="https://doi.org/10.5281/zenodo.3748491">https://doi.org/10.5281/zenodo.3748491</a></li> <li>CSEM data set: <a href="https://doi.org/10.5281/zenodo.1256786">https://doi.org/10.5281/zenodo.1256786</a></li> <li>Noise-free CSEM data set: <a href="https://doi.org/10.5281/zenodo.1807134">https://doi.org/10.5281/zenodo.1807134</a></li> </ul>
MultiCaRe: An open-source clinical case dataset for medical image classification and multimodal AI applications
<p>The dataset contains multi-modal data from over 70,000 open access and de-identified case reports, including metadata, clinical cases, image captions and more than 130,000 images. Images and clinical cases belong to different medical specialties, such as oncology, cardiology, surgery and pathology. The structure of the dataset allows to easily map images with their corresponding article metadata, clinical case, captions and image labels. Details of the data structure can be found in the file data_dictionary.csv.</p> <p>More than 90,000 patients and 280,000 medical doctors and researchers were involved in the creation of the articles included in this dataset. The citation data of each article can be found in the metadata.parquet file.</p> <p>Refer to the examples showcased in this <a href="https://github.com/mauro-nievoff/MultiCaRe_Dataset">GitHub repository</a> to understand how to optimize the use of this dataset.<br><br>The license of the dataset as a whole is CC BY-NC-SA. However, its individual contents may have less restrictive license types (CC BY, CC BY-NC, CC0). For instance, regarding image filess, 66K of them are CC BY, 32K are CC BY-NC-SA, 32K are CC BY-NC, and 20 of them are CC0.</p>
C2D2: An Open-Source, Pan-European, Harmonised Crop Development Database for Use in Regulatory Pesticide Exposure Modelling and Risk Assessment.
<p>There is a regulatory need for crop development dates to assess current default values used within chemical exposure assessments as well as to justify refinements within risk assessments. However, a readily available pan-European crop phenology database covering key FOCUS (FOrum for the Co-ordination of pesticide fate models and their USe) crops and scenarios to meet this need is not currently available. Therefore, we describe the development of a harmonised, pan-European, CropLife Europe Crop Development Database, C2D2, that is fully aligned with this regulatory requirement utilising efficacy trials data generated for regulatory submissions when registering plant protection products under Regulation (EU) 1107/2009. Evaluation of C2D2 against an independent dataset showed good agreement for equivalent time periods, crop growth stages and geographical regions. We illustrate how this database can be used to evaluate existing default crop development dates mandated by regulatory agencies for use within exposure assessments. Despite the large dataset compiled and the geographical coverage of C2D2, not all FOCUSsw/gw scenarios have sufficient data to facilitate comparison, with less significant scenarios, like FOCUSgw Porto, being under-represented. For those scenarios with sufficient data, clear differences between C2D2 and crop development dates assumed in the FOCUS modelling framework (using the AppDate tool) are often indicated over some/many growth stages suggesting that amendment of the existing representation of crop development within the risk assessment process may be required. C2D2 is freely available under a Creative Commons licence to facilitate innovation in exposure science to allow for more accurate and realistic risk assessment leading to enhanced crop and environmental protection.</p>
Scan4CFU: Low-cost, open-source bacterial colony tracking over large areas and extended incubation times
<p>A hallmark of bacterial populations cultured <em>in vitro</em> is their homogeneity of growth, where the majority of cells display identical growth rate, cell size and content. Recent insights, however, have revealed that even cells growing in exponential growth phase can be heterogeneous with respect to variables typically used to measure cell growth. Bacterial heterogeneity has important implications for how bacteria respond to environmental stresses, such as antibiotics. The phenomenon of antimicrobial persistence, for example, has been linked to a small subpopulation of cells that have entered into a state of dormancy where antibiotics are no longer effective. While methods have been developed for identifying individual non-growing cells in bacterial cultures, there has been less attention paid to how these cells may influence growth in colonies on a solid surface. In response, we have developed a low-cost, open-source platform to perform automated image capture and image analysis of bacterial colony growth on multiple nutrient agar plates simultaneously. The descriptions of the hardware and software are included, along with details about the temperature-controlled growth chamber, high-resolution scanner, and graphical interface to extract and plot the colony lag time and growth kinetics. Experiments were conducted using a wild type strain of <em>Escherichia coli </em>K12 to demonstrate the feasibility and operation of our setup. By automated tracking of bacterial growth kinetics in colonies, the system holds the potential to reveal new insights into understanding the impact of microbial heterogeneity on antibiotic resistance and persistence. </p>
Dataset of Open-Source Software Developers Labeled by their Experience Level and Associated with their Software Metrics
<p>This dataset contains 703 anonymized developers extracted from 17 open-source projects from GitHub. Projects were chosen because they use:</p> <ul> <li>the Java programming language</li> <li>the <a href="https://spring.io/projects/spring-framework">Spring framework</a></li> <li><a href="https://maven.apache.org/">Maven</a> / <a href="https://gradle.org/">Gradle</a> build tools</li> </ul> <p>For all these developers, 23 software metrics were calculated for each project to which they contribute. These metrics are either calculated by analyzing the source code or relative to project management metadata. Each of these developers then have been manually annotated. To do this, developers have been searched for in professionnal social media such as:</p> <ul> <li><a href="https://www.linkedin.com/">Linkedin</a></li> <li><a href="https://twitter.com/">Twitter</a></li> <li><a href="https://github.com/">Github</a></li> </ul> <p><strong>This dataset is published in the following journal article: </strong></p> <p><strong>Dataset of Open-Source Software Developers Labeled by their Experience Level in the Project and their Associated Software Metrics, Q. Perez, C. Urtado and </strong><strong>S. Vauttier, Data In Brief, </strong></p> <p><a href="https://www.sciencedirect.com/science/article/pii/S2352340922010459">https://www.sciencedirect.com/science/article/pii/S2352340922010459</a></p>
Open-source DGGS comparison data supplement
<p>A DGGS is a type of spatial reference system that partitions the globe into many individual, evenly spaced, and well-aligned cells to encode location. We calculated normalized area and compactness of cell geometries for 5 open-source DGGS implementations - Uber H3, Google S2, RiskAware OpenEAGGR, rHEALPix by Landcare Research New Zealand, HEALPix by NASA Jet Propulsion Labs, and DGGRID by Southern Oregon University - to evaluate their suitability for a global-level statistical data cube.</p> <p>This repository contains all generated data and statistics.</p> <ul> <li>EAGGR doesn't seem to have a predefined logic of hierarchical cell resolutions for ISEA3H</li> <li>EAGGR doesn't seem to have a region filling algorithm available, neither for ISEA4T nor ISEA3H</li> <li>rHEALPix is pure Python (with Numpy/Scipy support), but cell generation/conversion is slower than the other C/C++ based implementations</li> <li>DGGRID is a commandline tool and can predominantly only be used to generate a grid and fill with sampling data, the Python API is only a wrapper</li> <li>healpy is a Python package to handle pixelated data on the sphere. It is based on the Hierarchical Equal Area isoLatitude Pixelization (HEALPix) scheme and bundles the HEALPix C++ library.</li> </ul> <p>Kmoch et. al (2022). Area and Shape Distortions in Open-Source Discrete Global Grid Systems. <strong><em>Big Earth Data</em></strong></p>
Dataset - DeepWealth: A Generalizable Open-Source Deep Learning Framework using Satellite Images for Well-Being Estimation
<p>This dataset encapsulates the Checkpoints obtained during the training process of the Deep Learning model, which can be used for new estimations.</p> <p>The aim of the DeepWealth package is to provide a generalizable Deep Learning framework for the use of remote sensing in poverty estimation. The combination of Deep Learning and Earth Observation data is increasingly being used to estimate socioeconomic conditions at regional and global scales. The proposed framework aligns with the Sustainable Development Goal SDG1 of ending poverty. The framework provides open-source data, code, and training models (checkpoints) for reproducibility and replicability.</p> <ul> <li>The source code can be found in <a href="https://github.com/PARSECworld/DeepWealth" target="_blank" rel="noopener">https://github.com/PARSECworld/DeepWealth</a></li> <li>The metadata from source code can be found in <a href="https://github.com/PARSECworld/DeepWealth/blob/main/metadata.pdf" target="_blank" rel="noopener">https://github.com/PARSECworld/DeepWealth/blob/main/metadata.pdf</a></li> <li>The paper describing the development of this framework can be found at: Ben Abbes, A., Machicao, J., Corrêa, P. L. P., Specht, A., Devillers, R., Ometto, J. P., Kondo, Y., & Mouillot, D. (2024). DeepWealth: A generalizable open-source deep learning framework using satellite images for well-being estimation. <em>SoftwareX</em>, 27, 101785. <a href="https://doi.org/10.1016/j.softx.2024.101785">https://doi.org/10.1016/j.softx.2024.101785</a> </li> </ul>
Zambezi dataset to "WHAT-IF: an open-source decision support tool for water infrastructure investment planning within the Water-Energy-Food-Climate Nexus"
<p>This is the dataset used in the HESS publication "<a href="https://www.hydrol-earth-syst-sci-discuss.net/hess-2019-167/">WHAT-IF: an open-source decision support tool for water infrastructure investment planning within the Water-Energy-Food-Climate Nexus</a>"</p> <p>The dataset describes the water-energy-food nexus of the Zambezi River Basin used as input to the <a href="https://github.com/RaphaelPB/WHAT-IF">WHAT-IF model</a>.</p> <p>The file Data_Organization.pdf, summarizes the available data. For more info look at the <a href="https://www.hydrol-earth-syst-sci-discuss.net/hess-2019-167/">publication</a> and/or <a href="https://github.com/RaphaelPB/WHAT-IF">Github</a>.</p>
Accompanying data for the open-source book Modeling of Hydrological Systems in Semi-Arid Central Asia
<p>This data set is used to reproduce examples in the open-source book <a href="https://hydrosolutions.github.io/caham_book/">"Modeling of Hydrological Systems in Semi-Arid Central Asia"</a> which is part of a free course on hydrological modeling in Central Asia. The course teaches how to use publicly available data to implement a hydrological model for climate impact studies (Marti et al., 2023). </p> <p>To use the data set to reproduce the examples in the book: Download the book from https://doi.org/10.5281/zenodo.6350042 and this data set to the same hierarchical level in your file system: </p> <p>|- caham_book<br> |- caham_data<br> |- AmuDarya<br> |- central_asia_domain<br> |- student_case_study_basins<br> |- SyrDarya</p> <p>You will need a working installation of R (https://www.r-project.org/) and a GUI (e.g. Posit, formerly RStudio https://posit.co/) to reproduce the scripted examples in the book. Once your software is set up, you can proceed to run the examples. </p> <p> </p>
An Open-Source Automatic Survey of Green Roofs in London using Segmentation of Aerial Imagery: Dataset
<p>This archive contains code and data to go with the paper <em>*An Open-Source Automatic Survey of Green Roofs in London using Segmentation of Aerial Imagery*</em>.</p> <p> </p> <p>This archive contains geospatial data, as well as the code used to generate the geospatial data.</p> <p>The geospatial data consists of georeferenced polygons identifying areas which are covered by green roofs in London (GBR) generated from 2019 aerial imagery.</p> <p>The data is described in detail in the manuscript <em>*An Open-Source Automatic Survey of Green Roofs in London using Segmentation of Aerial Imagery*</em>. See abstract below.</p> <p> </p> <p>GeoJSON format:</p> <p>GeoJSON is a format for encoding geospatial data, see https://geojson.org/.</p> <p>GeoJSON can be read using GIS programs including ArcGIS, QGIS, OGR.</p> <p> </p> <p>Contents:</p> <p>`geospatial_data/buffered_polygons_2021.zip` a zip archive containing a geojson file. It is the estimated locations of green roofs in London in 2021 and is the main result, which can be opened in any GIS program after being unzipped.</p> <p>`geospatial_data/buffered_polygons_2019.zip` a zip archive containing a geojson file. It is the estimated locations of green roofs in London in 2019 and is a secondary result, which can be opened in any GIS program after being unzipped. The predictions were made with the same model as the 2021 results.</p> <p>`geospatial_data/labelled_area.zip` a zip archive containing a geojson file. Identifies the area which was hand-labelled.</p> <p>`geospatial_data/manual_2021.zip` a zip archive containing a geojson file. Manually labelled green roof from 2021 imagery.</p> <p>`geospatial_data/manual_2019.zip` a zip archive containing a geojson file. Manually labelled green roof from 2019 imagery.</p> <p>`segmentation_code` contains the code used to produce the segmentation from the aerial imagery.</p> <p>`analysis_code` contains the code used to produce the plots and tables for the paper.</p> <p> </p> <p>Imagery availability:</p> <p>Unfortunately the aerial imagery and building footprint data cannot be shared directly, as you will require the proper license. Both can be found at [Digimap](https://digimap.edina.ac.uk) provided your institution has the license.</p> <p> </p> <p>Abstract:</p> <p>Green roofs can mitigate heat, increase biodiversity, and attenuate storm water, giving some of the benefits of natural vegetation in an urban context where ground space is scarce. To guide the design of more sustainable and climate resilient buildings and neighbourhoods, there is a need to assess the existing status of green roof coverage and explore the potential for future implementation. Therefore, accurate information on the prevalence and characteristics of existing green roofs is needed, but this information is currently lacking. Segmentation algorithms have been used widely to identify buildings and land cover in aerial imagery. Using a machine-learning algorithm based on U-Net to segment aerial imagery, we surveyed the area and coverage of green roofs in London, producing a geospatial dataset \cite[]{simpson_charles_2022_6861929}. We estimate that there was 0.23 km^2 of green roof in the Central Activities Zone (CAZ) of London, (1.07 km^2) in Inner London, and (1.89 km^2) in Greater London in the year 2021. This corresponds to 2.0% of the total building footprint area in the CAZ, and 1.3% in Inner London. There is a relatively higher concentration of green roofs in the City of London, covering 3.9% of the total building footprint area. Test set accuracy was 0.99, with an f-score of 0.58. When tested against imagery and labels from a different year (2019), the model performed just as well as a model trained on the imagery and labels from that year, showing that the model generalised well between different imagery. We improve on previous studies by including more negative examples in the training data, and by requiring coincidence between vector building footprints and green roof patches. We experimented with different data augmentation methods, and found a small improvement in performance when applying random elastic deformations, colour shifts, gamma adjustments, and rotations to the imagery. The survey covers 1558 km^2 of Greater London, making this the largest open automatic survey of green roofs in any city. The geospatial dataset is at the single-building level, providing a higher level of detail over the larger area compared to what was already available. This dataset will enable future work exploring the potential of green roofs in London and on urban climate modelling.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.