Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
867
datasets available to search
ShareScore release 0.9.0
Dataset results
867 results for “repositories”
alexlwhite/WhiteBoyntonYeatman2019_Repository: Public repository for White, Boynton & Yeatman (2019)
<p>This is the public repository of data and code for the study described in our 2019 Cortex article, "The link between reading ability and visual spatial attention across development."</p>
Final Pool and Online Repository
<p>This file includes all of the classification details of <strong><em>"Quality and Success in Open Source Software: A Systematic Mapping" study.</em></strong></p>
Data Repository for: SOCIO-ENVIRONMENTAL IMPACTS OF OIL PALM CONTRACT FARMING SCHEMES IN THE BRAZILIAN AMAZON
<p>Contract farming is arguably a pro-poor strategy to promote rural development and minimize the social impacts of large-scale agricultural expansion. Yet, little attention has been paid to its environmental impacts, particularly in tropical landscapes. This article fills this gap by linking social and environmental analysis of oil palm contract farming in the Brazilian Amazon. The analysis presented used a mix of quantitative and qualitative methods, including household surveys, remote sensing techniques, and in-depth interviews, to assess whether the scheme managed to avoid deforestation and to contribute to livelihood improvements. The results show that the Brazilian model managed to prevent the deforestation of primary forests, but achieved limited and differentiated livelihood results. The analysis suggests that the Brazilian model is more likely to work for households with an agricultural vocation and a commercial spirit in areas with an abundant availability of degraded lands, but can hardly be a solution in frontier areas or for subsistence or more dependent households. The chapter concludes with some reflections on how contract farming schemes should be designed in order to maximize livelihood gains and minimize negative environmental impacts.</p>
Repository: Potential for Photosynthesis on Mars within snow and ice
<p>This repository contains:</p> <p>1. Modeled Spectral Irradiances (W m-2 micron-1) within Vertically Inhomogeneous Glacier Ice from Khuller, Warren, Christensen & Clow (2024)</p> <p>a) Fig_1_12cm_model: Modeled Spectral Irradiance at 12 cm <br>b) Fig_1_36cm_model: Modeled Spectral Irradiance at 36 cm <br>c) Fig_1_58cm_model: Modeled Spectral Irradiance at 58 cm <br>d) Fig_1_77cm_model: Modeled Spectral Irradiance at 77 cm </p> <p>2. Modeled Spectral Actinic Flux (W m-2 micron-1) from Khuller, Warren, Christensen & Clow (2024)</p> <p>a) Clean Snow/Firn/Ice (without dust) at 33 S latitude<br> i) Fig_2a: Pure snow with 0.5 mm grain size<br> ii) Fig_2b: Pure firn with 2.5 mm grain size<br> iii) Fig_2c: Pure glacier ice with 14 mm grain size</p> <p>b) Clean Snow/Firn/Ice (without dust) at 54 N latitude<br> i) Ext_Fig_2a: Pure snow with 0.5 mm grain size<br> ii) Ext_Fig_2b: Pure firn with 2.5 mm grain size<br> iii) Ext_Fig_2c: Pure glacier ice with 14 mm grain size </p> <p>c) Dusty Snow/Firn/Ice (with 0.01% dust by mass) at 33 S latitude<br> i) Fig_2d: Dusty snow with 0.5 mm grain size<br> ii) Fig_2e: Dusty firn with 2.5 mm grain size<br> iii) Fig_2f: Dusty glacier ice with 14 mm grain size</p> <p>d) Dusty Snow/Firn/Ice (with 0.01% dust by mass) at 54 N latitude<br> i) Ext_Fig_2d: Dusty snow with 0.5 mm grain size<br> ii) Ext_Fig_2e: Dusty firn with 2.5 mm grain size<br> iii) Ext_Fig_2f: Dusty glacier ice with 14 mm grain size</p> <p>e) Dusty Snow/Firn/Ice (with 0.1% dust by mass) at 33 S latitude<br> i) Fig_2g: Dusty snow with 0.5 mm grain size<br> ii) Fig_2h: Dusty firn with 2.5 mm grain size<br> iii) Fig_2i: Dusty glacier ice with 14 mm grain size</p> <p>f) Dusty Snow/Firn/Ice (with 0.1% dust by mass) at 54 N latitude<br> i) Ext_Fig_2g: Dusty snow with 0.5 mm grain size<br> ii) Ext_Fig_2h: Dusty firn with 2.5 mm grain size<br> iii) Ext_Fig_2i: Dusty glacier ice with 14 mm grain size </p> <p>3. Modeled Depths for DNA Damage Limit, PAR Upper Limit, and PAR Lower Limit (meters) from Khuller, Warren, Christensen & Clow (2024)</p> <p>a) Sensitivity to dust content<br>FILE FORMAT: Dust Content (ppmw), DNA Damage Limit Depth (m), PAR Lower Limit Depth (m), PAR Upper Limit Depth (m)<br> i) final_Depths_south_dust_sensitivity: sensitivity to dust content for the martian southern hemisphere<br> ii) final_Depths_north_dust_sensitivity: sensitivity to dust content for the martian northern hemisphere</p> <p>b) Sensitivity to ice grain radius<br>FILE FORMAT: Ice Grain Radius (micron), DNA Damage Limit Depth (m), PAR Lower Limit Depth (m), PAR Upper Limit Depth (m)<br> i) final_Depths_south_radius_sensitivity: sensitivity to ice grain radius for the martian southern hemisphere<br> ii) final_Depths_north_radius_sensitivity: sensitivity to ice grain radius for the martian northern hemisphere</p> <p>c) Sensitivity to latitude<br>FILE FORMAT: Latitude (degrees), DNA Damage Limit Depth (m), PAR Lower Limit Depth (m), PAR Upper Limit Depth (m)<br> i) final_Depths_south_latitude_sensitivity: sensitivity to latitude for the martian southern hemisphere<br> ii) final_Depths_north_latitude_sensitivity: sensitivity to latitude for the martian northern hemisphere</p> <p>d) Sensitivity to solar zenith angle<br>FILE FORMAT: Solar Zenith Angle (degrees), DNA Damage Limit Depth (m), PAR Lower Limit Depth (m), PAR Upper Limit Depth (m)<br> i) final_Depths_south_zenith_sensitivity: sensitivity to solar zenith angle for the martian southern hemisphere<br> ii) final_Depths_north_zenith_sensitivity: sensitivity to solar zenith angle for the martian northern hemisphere</p> <p>4. Wavelengths used for files listed in 1 from Khuller, Warren, Christensen & Clow (2024)<br>wavelengths_greenland: wavelength in microns</p> <p>5. Wavelengths used for files listed in 2 and 3 from Khuller, Warren, Christensen & Clow (2024)<br>wavelengths: wavelength in microns</p> <p>6. Depths used for files listed in 2 from Khuller, Warren, Christensen & Clow (2024)<br>depths: depths in meters</p> <p>7. Normalized DNA spectrum used in Khuller, Warren, Christensen & Clow (2024)<br>norm_dna_spectrum: normalized DNA spectrum</p> <p> </p>
Repository for: Combinatorial Wnt signaling landscape during brachiopod anteroposterior patterning
<p>This repository contains the data and analyses for the manuscript:</p> <p>Vellutini, B. C., Martín-Durán, J. M., Børve, A. & Hejnol, A. <strong>Combinatorial Wnt signaling landscape during brachiopod anteroposterior patterning.</strong> BMC Biol. 22, 1–23 (2024). <a href="https://doi.org/10.1186/s12915-024-01988-w" rel="nofollow">https://doi.org/10.1186/s12915-024-01988-w</a></p> <p>The source is maintained at <a href="https://github.com/bruvellu/terebratalia-wnts">https://github.com/bruvellu/terebratalia-wnts</a>.</p>
Repository for: "Using automatic calibration to improve the physics behind complex numerical models: An example from a 3D lake model"
<p>Set of numerical experiments supporting the paper entitled "Using automatic calibration to improve the physics behind complex numerical models: An example from a 3D lake model" by Marina Amadori, Abolfazl Irani Rahaghi, Damien Bouffard and Marco Toffolon. Submitted to GMD. </p> <p>The folder contains: </p> <p>simulations: DYNO-PODS + Delft3D experiments on Lake Morat. See https://github.com/louisXW/DYNO-pods for more insights on DYNO-PODS and instructions for installation.</p> <p>scripts: extraction and plotting scripts</p> <p>source_code: modified Delft3D src as available at: https://github.com/eawag-surface-waters-research/Delft3D/tree/d3d4/research/surface_heat_transfer</p>
Grass Phylogeny Working Group III: data repository
<p><strong>Grass Phylogeny Working Group III: data repository</strong></p> <p>Phylogenetic analyses of the grass family (Poaceae) using nuclear and plastid data. The data set includes 1153 accessions corresponding to 1133 accepted species. Genomic data was obtained from different sources including target capture, shotgun, transcriptomes and annotated genomes. Nuclear markers (Angiosperm353 gene set) were assembled from short read data using HybPiper or a custom assembly pipeline optimized for low coverage shotgun data. Plastid genes were either retrieved from published plastome sequences or assembled here using getOrganelle. This data set also includes the results of a gene tree-species tree reconciliation analysis using GeneRax.</p> <p> </p> <p>Contact persons:</p> <p>Matheus E. Bianconi (matheus-enrique.bianconi@univ-tlse3.fr), Jan Hackel (jan.hackel@uni-marburg.de), Maria S. Vorontsova (m.vorontsova@kew.org)</p> <p> </p> <p>Content description</p> <p><strong>1. Metadata</strong></p> <ul> <li><code>gpwgIII_samples_metadata_taxonomy.tsv</code></li> </ul> <p>Tab-separated file with details for all 1,702 accessions used in this study. Columns: analysis_ID - ID in nuclear analyses; analysis_ID_plastome - ID in plastome analyses; acc_species - accepted species name; acc_species_author - taxonomic species authority; acc_genus - accepted genus name; acc_genus_author - taxomomic genus authority; publication - associated prior publication; data type - type of sequence data; isolate - laboratory isolate ID; voucher_ID - herbarium voucher ID; germplasm_ID - germplasm collection ID; repo_accession - accession number in public repository; plastome_accession - accession number of assembled plastome sequence; removed_nuclear - reason for removal from nuclear tree, if applicable; removed_plastome - reason for removal from plastome tree, if applicable; soreng2022_genus - genus name in Soreng et al. 2022, https://doi.org/10.1111/jse.12847; subtribe, tribe, subfamily, major.clade - classification according to Soreng et al. 2022.</p> <p><strong>2. Nuclear data</strong></p> <p><em>- Dataset1 ("main")</em><br>Number of samples: 1153<br>Number of genes: 331<br>Alignment trimming threshold: gt = 0.1 (removed sites > 90% missing data)<br>Genes per sample: > 166</p> <p><em>- Dataset2 ("strict trimming")</em><br>Number of samples: 1153<br>Number of genes: 315<br>Alignment trimming threshold: gt = 0.5 (removed sites > 50% missing data)<br>Genes per sample: > 158</p> <p><em>- Dataset3 (dataset 1 without shotgun samples)</em><br>Number of samples: 841<br>Number of genes: 331<br>Alignment trimming threshold: gt = 0.1 (removed sites > 90% missing data)<br>Genes per sample: > 166</p> <p><strong>2.1. Raw sequences</strong></p> <p>Raw Ang353 sequence assemblies for all samples (pre-trimming and filtering)</p> <ul> <li><code>raw_Ang353_sequences.zip</code></li> </ul> <p><strong>2.2 Nuclear gene alignments</strong></p> <p>Trimmed alignments from datasets 1, 2 and 3.</p> <ul> <li><code>alignments_dataset1_main_final.zip</code></li> <li><code>alignments_dataset2_strict_trimming_final.zip</code></li> <li><code>alignments_dataset3_no_shotgun_final.zip</code></li> </ul> <p><strong>2.3. Nuclear gene trees</strong><br>Gene trees inferred using RAxML (GTRCAT, 100 bootstraps) for the alignments from datasets 1, 2 and 3.</p> <ul> <li><code>gene_trees_dataset1_main_final.zip</code></li> <li><code>gene_trees_dataset2_strict_trimming_final.zip</code></li> <li><code>gene_trees_dataset3_no_shotgun_final.zip</code></li> </ul> <p><strong>2.4. Multigene species trees</strong><br>Multigene species trees obtained using Astral-Pro3 from gene trees for datasets 1, 2 and 3. </p> <ul> <li><code>astralpro_trees.zip</code>, which includes: <ul> <li>trees_Ang353_grasses_dataset1_main_gtrcat.astralpro</li> <li>trees_Ang353_grasses_dataset2_strict_trimming_gtrcat.astralpro</li> <li>trees_Ang353_grasses_dataset3_no_shotgun_gtrcat.astralpro</li> </ul> </li> </ul> <p><strong>3. Gene tree–species tree reconciliation</strong></p> <ul> <li><code>generax.zip</code></li> </ul> <p>Compressed zip archive with input files and results, including log files, of the GeneRax reconciliation analysis. One subfolder for each of the four analyses run: "all_tribes", "Andropogoneae", "Bambusoideae", "Triticeae".</p> <ul> <li><code>transfers_reconciliation_analyses.zip</code>, which includes: <ul> <li>transfers_all_all_tribes.tsv: Tab-separated file with all transfers inferred with the tribe-level Poaceae reconciliation analysis. Each line represents one transfer inferred.</li> <li>transfers_all_Andropogoneae.tsv: Tab-separated file with all transfers inferred with the Andropogoneae reconciliation analysis. Each line represents one transfer inferred.</li> <li>transfers_all_Bambusoideae.tsv: Tab-separated file with all transfers inferred with the Bambusoideae reconciliation analysis. Each line represents one transfer inferred.</li> <li>transfers_all_Triticeae.tsv: Tab-separated file with all transfers inferred with the Triticeae reconciliation analysis. Each line represents one transfer inferred.</li> <li>transfers_counts_all_tribes.tsv: Tab-separated file with aggregated transfer counts, in both directions for each reticulate connection, from the tribe-level Poaceae reconciliation analysis.</li> <li>transfers_counts_Andropogoneae.tsv: Tab-separated file with aggregated transfer counts, in both directions for each reticulate connection, from the Andropogoneae reconciliation analysis.</li> <li>transfers_counts_Bambusoideae.tsv: Tab-separated file with aggregated transfer counts, in both directions for each reticulate connection, from the Bambusoideae reconciliation analysis.</li> <li>transfers_counts_Triticeae.tsv: Tab-separated file with aggregated transfer counts, in both directions for each reticulate connection, from the Triticeae reconciliation analysis.</li> </ul> </li> </ul> <p><strong>4. Plastome data</strong></p> <p>Alignment and phylogenetic tree from plastome data.</p> <ul> <li><code>plastome_files.zip</code>, which includes <ul> <li>reduced_plastome_concat_CDS_trnLtrnF_trimmed.fna-out.fas: FASTA file with the final, concatenated DNA alignment of 71 plastome regions for 910 accessions, after data filtering.</li> <li>partitions.txt: Text file with positions of the 71 plastome regions in the concatenated alignment.</li> <li>plastome_concat_CDS_trnLtrnF_trimmed_TBE.raxml.support: Plastome tree with Transfer Bootstrap Expectation values as node labels.</li> <li>RAxML_bipartitions.plastome_concat_CDS_trnLtrnF_trimmed: Maximum likelihood plastome tree inferred with RAxML, with Felsenstein bootstrap values as node labels.</li> <li>RAxML_bootstrap.plastome_concat_CDS_trnLtrnF_trimmed: 100 rapid bootstrap pseudoreplicate plastome trees inferred with RAxML.</li> <li>RAxML_info.plastome_concat_CDS_trnLtrnF_trimmed: RAxML analysis log file.</li> <li>nuc_plastome_matching_tips.tab: Tab-separated file with accessions matched in nuclear-plastome comparison.</li> </ul> </li> </ul> <p><strong>5. Poaceae-specific reference Ang353 dataset</strong><br>Reference sequence dataset used for the assembly of Ang353 sequences in this study.</p> <ul> <li><code>target_Ang353_sequences_grasses.zip</code></li> </ul> <p><strong>6. Shotgun assembly script</strong></p> <p>Custom script used for the assembly of Ang353 sequences from shotgun data</p> <ul> <li><code>shotgun_assembler_script.zip</code>, which includes: <ul> <li>shotgun_assembler_Ang353_sequences.sh: script for assembly of short reads from shotgun data</li> <li>template_manifest_file.tsv: TAB-separated file to specify sample names and location of short read files (required by the assembly script)</li> <li>list_Ang353_genes_orthofinder.txt: list of Ang353 gene identifiers (required by the assembly script)</li> </ul> </li> </ul> <p><strong>7. Quartet metrics script</strong></p> <p>R script to calculate the Quartet Concordance (QC) and Quartet Differential (QD) metrics from the gene tree frequencies/proportions for each quartet at a branch, following Pease et al. 2018 (American Journal of Botany, <span><a href="https://doi.org/10.1002/ajb2.1016" target="_blank" rel="nofollow noopener noreferrer">https://doi.org/10.1002/ajb2.1016</a></span>).</p> <ul> <li><code>quartet_metrics.R</code></li> </ul>
Compilation of data collected in surveys on the WissKI-based 3D Repository with DFG 3D-Viewer project partners and architecture students from the Warsaw University of Technology and the Technical University of Łódź
<p>The dataset contain the compiliation of responses from users of WissKI-based 3D Repository (https://3d-repository.hs-mainz.de/)., which is the open platform for deposit of 3D models of cultural heritage. The beta version of the WissKI 3D Repository, initiated in June 2022, has been subjected to evaluation by two primary target groups since its launch. The initial group, composed of students in the cultural heritage domain, was tasked with showcasing the importance of documenting and publishing 3D models of digital reconstructions. The survey with sutdents was conducted for three different classes: </p> <p>1) In summer 2022 with bachelor architectrue students at Warsaw University of Technology during seminar of choice regarding digital reconstruction of wooden synagogues;</p> <p>2) In summer 2023 with bachelor architectrue students at Warsaw University of Technology, and master students from Technology University of Łódź during seminar of choice regarding digital reconstruction of wooden synagogues;</p> <p>3) In autumn 2023 during international workshop about digital 3D heritage of CoVHer project with studnets of architecture from Warsaw Univeristy of Technology, Alma Mater Studiorum – Universita di Bologna, Facoltà di Architettura di Porto and Hochschule Mainz - University of Applied Sciences, as well as archaeology studnets from Universitat Autònoma de Barcelona.</p> <p>The second group, comprising digital 3D cultural heritage professionals, predominantly focused on archiving digital assets. Participants were project partners of DFG 3D Viewer project, which were professionals from the Institute of Archaeology at University Cologne, the Institute of Art History at the Ludwig-Maximilians-Universität Munich, the Architecture, Civil Engineering and Urban Planning Department of BTU Cottbus Senftenberg, and the Detushce Museum. They were asked for evaluaton of system after three differetn stages of work: at the begging wihtout any introduction to the system, after proivision of intorudctionary materilas and finally at the end of work.</p> <p>All participants were requested to report their experiences across four categories: metadata form, 3D viewer, provided guidelines, and overall experience. A 5-point rating scale was employed to assess specific issues, with 1 being the most negative and 5 being the most positive. The form length question was an exception, where a median value of 3 was considered ideal, and extreme values indicated either excessive length or brevity.</p>
Experimental Repository for "Certifying Without Loss of Generality Reasoning in Solution-Improving Maximum Satisfiability"
<h2>Experimental repository for "Certifying Without Loss of Generality Reasoning in Solution-Improving Maximum Satisfiability".</h2> <h3>The directory is structured as follows:</h3> <ul> <li><code>data</code>: Data that has been processed into CSVs and the scripts to analyse the experiments.</li> <li><code>plots</code>: The plots generated from the data that are used in the paper.</li> <li><code>raw_data</code>: The raw data logs for the experiments and scripts to extract relevant data from the logs.</li> <li><code>source_code</code>: Source code for the checker <code>VeriPB</code> and the MaxSAT solver Pacose in the different variants used for the experiments. The Pacose version in <code>PacoseMaxSATSolver-baseline</code> is Pacose without proof logging, the version in <code>PacoseMaxSATSolver-certified</code> is Pacose with proof logging. Proof logging using only assumptions for the coarse convergence can be enabled via the option <code>--WithAssumptions</code>.</li> </ul> <h3>Install Requirements</h3> <p>To install and run the MaxSAT solver Pacose and the pseudo-Boolean proof checker VeriPB your need to have the following components installed:</p> <ul> <li>Python 3.6.9 or higher with pip and setuptools installed</li> <li>g++ 7.5.0 or higher</li> <li>libgmp</li> </ul> <p>These can be installed in Ubuntu / Debian via</p> <pre><code>sudo apt-get update && apt-get install \ python3 \ python3-pip \ python3-dev \ g++ \ libgmp-dev pip3 install --user \ setuptools</code></pre> <h3>How to Run?</h3> <p>The MaxSAT solver Pacose can be compiled using the install script in <code>PacoseMaxSATSolver-certified</code>:</p> <p><code>./install</code></p> <p>To run Pacose with proof logging, where the proof should be written to <code>proof.pbp</code>, run the following command in the <code>PacoseMaxSATSolver-certified</code> directory:</p> <p><code>./bin/Pacose --proofFile proof.pbp instance.wcnf</code></p> <p>The proof can be checked with VeriPB. To compile VeriPB run the following in the <code>VeriPB</code> directory:</p> <p><code>pip install .</code></p> <p>To check the proof with VeriPB, run the following inside the <code>PacoseMaxSATSolver-certified</code> directory:</p> <p><code>veripb --wcnf instance.wcnf proof.pbp</code></p>
AgroRadarEval - Data and Code Repository.
<p>Contains the data, code, and supplementary materials associated with the article "Data-driven RD&I management for societal impacts through evaluation results: introducing and applying AgroRadarEval". AgroRadarEval aims to support leaders and managers in agricultural RD&I by reflecting on organizational capacities, culture, collaborations, processes, and communications underlying the use of evaluation results. The approach incorporates principles from Responsible Research and Innovation (RRI) and Responsible Research Assessment (RRA) to ensure that evaluation processes are inclusive, transparent, and aimed at societal impact.</p>
Linked collectors and determiners for: The ITEM fungal repository dataset (ISPA-ITEM-02).
Natural history specimen data linked to collectors and determiners held within, "The ITEM fungal repository dataset (ISPA-ITEM-02)". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/e5347845-c5bd-4cff-9a5d-83ba9a7b1978">https://bionomia.net/dataset/e5347845-c5bd-4cff-9a5d-83ba9a7b1978</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/e5347845-c5bd-4cff-9a5d-83ba9a7b1978">https://gbif.org/dataset/e5347845-c5bd-4cff-9a5d-83ba9a7b1978</a>. Formatted as a Frictionless Data package.
Linked collectors and determiners for: The ITEM fungal repository dataset (ISPA-ITEM-01).
Natural history specimen data linked to collectors and determiners held within, "The ITEM fungal repository dataset (ISPA-ITEM-01)". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/18abd151-94fd-40a9-b9c4-3bc7e3af7af1">https://bionomia.net/dataset/18abd151-94fd-40a9-b9c4-3bc7e3af7af1</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/18abd151-94fd-40a9-b9c4-3bc7e3af7af1">https://gbif.org/dataset/18abd151-94fd-40a9-b9c4-3bc7e3af7af1</a>. Formatted as a Frictionless Data package.
EPSRC HEED Data Repository: Nepal Household Appliance Survey
<p>The dataset deposited here was prepared under the EPSRC-funded <a href="http://heed-refugee.coventry.ac.uk/">Humanitarian Engineering and Energy for Displacement</a> research project (EP/P029531/1). The project aimed to understand the energy needs of displaced communities, create an evidence base on the usage of different energy interventions and provide recommendations for improved design of future energy interventions to better meet the needs of people. </p> <p>As part of the project, three Appliance surveys were conducted in the Uttargaya settlement in Nepal. Appliance surveys are designed to assess the energy needs of a community based on the devices they use. The surveys span three instances across 18 months, starting in October 2018 and ending in April 2020.</p> <p>The survey splits the participants into four categories, organised into sheets, based on the type of metering participants have: 'bulk meter'; 'sub meter'; 'do not possess meter' and 'do not have electrical connection'. The survey anonymises the name of participants and assigns them a unique id as a household number. Information is recorded on the gender of the household owner, the number of people in the household, the type of their electricity connection and the payment type for the electricity connection. The survey collects information on how many of the following appliances have: Electric Bulb; Mobile charger; Refrigerator; Television; Electric Radio; Table Fan; Electric Iron.</p>
GitHub Issue Dataset From Top Repositories of Top Languages
<p>GitHub issue dataset from top 200 most popular repositories associated with top 55 programming languages. The language ranking used for this work is available at: <a href="https://spectrum.ieee.org/static/interactive-the-top-programming-languages-2020">https://spectrum.ieee.org/static/interactive-the-top-programming-languages-2020</a></p> <p>Original work utilizes this dataset to classify issue reports into respective categories: Available at <a href="https://github.com/ansnadeem/aic">https://github.com/ansnadeem/aic</a></p> <p>The original work also appeared in ISSRE'21 titled '<strong>Automatic Issue Classifier: A Transfer Learning Framework for Classifying Issue Reports</strong>'. Please consider citing our work if you use this dataset.</p>
Repository Analytics and Metrics Portal (RAMP) 2021 data
<p>The Repository Analytics and Metrics Portal (RAMP) is a web service that aggregates use and performance use data of institutional repositories. The data are a subset of data from RAMP, the Repository Analytics and Metrics Portal (<a href="http://ramp.montana.edu/">http://rampanalytics.org</a>), consisting of data from all participating repositories for the calendar year 2021. For a description of the data collection, processing, and output methods, please see the "methods" section below.</p> <p>The record will be revised periodically to make new data available through the remainder of 2021.</p>
Data Repository - Thermal-electrochemical parametrisation of a lithium-ion battery: mapping Li concentration and temperature dependencies
<p>Datasets from "Thermal-electrochemical parametrisation of a lithium-ion battery: mapping Li concentration and temperature dependencies" - Journal of Electrochemical Society.</p> <p>This repository contains parameter values for the electrode solid-state diffusivity, entropic term, exchange current density, electronic conductivity, specific heat capacity, and thermal conductivity.</p>
New estimation of the NOx snow-source on the Antarctic Plateau - repository dataset
<p>Notebook and data set used to present the results.</p> <p>For more information, please, do not hesitate to contact the corresponding author to this study.</p>
Binary Classification as a Phase Separation Process (data repository)
<p><strong>For version 0.0.2 (from 2021) see below:</strong></p> <p>This is a data repository for the paper "Binary classification as a phase separation process", by Rafael Monteiro.</p> <ul> <li>Website with description of this project: <a href="https://rafael-a-monteiro-math.github.io/Binary_classification_phase_separation/index.html">https://rafael-a-monteiro-math.github.io/Binary_classification_phase_separation/index.html</a></li> <li>Github: <a href="https://github.com/rafael-a-monteiro-math/Binary_classification_phase_separation">https://github.com/rafael-a-monteiro-math/Binary_classification_phase_separation</a></li> </ul> <p>This is a second version, which I wrote using tensorflow. It is much smaller (5 Gb when decompressed), a remarkable improvement when compared to the more than 100 Gb of the previous version).</p> <p>The new files are </p> <ul> <li>PSBC_BCs.tar.gz</li> <li>PSBC_classifier_PCA.tar.gz</li> <li>PSBC_dataset.tar.gz</li> <li>PSBC_libs_grids_statistics.tar.gz</li> <li> PSBC_notebooks.tar.gz</li> </ul> <p>Their content is explained in the file README_v2.pdf</p> <p><strong>UPDATE: <a href="https://drive.google.com/drive/folders/18l_92HuHDWJDkZnvXRuyGedcyC_3YZ2M?usp=sharing">a Google Colab folder is also available</a>. You can also find all the data and libraries there, unpacked.</strong></p> <p>For usage, see the Git-hub. </p> <blockquote> <p><strong>NOTE)</strong> I will keep the content for the previous version available in my Github as well. It is still a "nice exercise" to do all that is done in this new version in numpy, as done there. <strong><em>(Or, I should say, they should be studied as a cautionary tale of what to avoid.)</em></strong></p> </blockquote> <p> </p> <p><strong>For version 0.0.1 (from 2020) see below:</strong></p> <p>This is a data repository for the paper "Binary classification as a phase separation process", by Rafael Monteiro.</p> <ul> <li>Website with description of this project: <a href="https://rafael-a-monteiro-math.github.io/Binary_classification_phase_separation/index.html">https://rafael-a-monteiro-math.github.io/Binary_classification_phase_separation/index.html</a></li> <li>Github: <a href="https://github.com/rafael-a-monteiro-math/Binary_classification_phase_separation">https://github.com/rafael-a-monteiro-math/Binary_classification_phase_separation</a></li> </ul> <p>Therein you will find</p> <ul> <li>Examples</li> <li>1D toy model examples</li> <li>Computational statistics</li> <li>Several trained PSBC on MNIST dataset, with different parameter configurations</li> <li>Extra simulations, investigating normalization properties, low dimensional models that fail due to "too much" model compression, and comparison among ANNs, KNNs, and the PSBC in 1D</li> </ul> <p>If you want to know</p> <ol> <li>how to read the data</li> <li>how to access computational statistics, raw data, and examples</li> <li>how to use the data stored in this data repository</li> </ol> <p>see the guide README.pdf on GitHub page at <a href="https://github.com/rafael-a-monteiro-math/Binary_Classification_Phase_Separation">Binary_Classification_Phase_Separation</a>, where a script that downloads (and organizes) all this data is also available ("download_PSBC.sh).</p> <p>I did not include a copy of the train-test set (0-1dubset of the MNIST database) in every folder with simulations. But you can find a copy of the normalized dataset in the tar ball "PSBC_Examples.tar.gz" as</p> <p>data_test_normalized_MNIST.csv and data_train_normalized_MNIST.csv.</p> <p> </p>
Madagascar's extraordinary biodiversity: a data repository
<p>Data repository for the two sister reviews of Madagascar's biodiversity:</p> <ul> <li>Antonelli et al.: "Madagascar's extraordinary biodiversity: Evolution, distribution, and use", Science 378 (6623): eabf0869 – <a href="https://doi.org/10.1126/science.abf0869">https://doi.org/10.1126/science.abf0869</a></li> <li>Ralimanana et al.: "Madagascar's extraordinary biodiversity: Threats and opportunities", Science 378 (6623): eadf1466 – <a href="https://doi.org/10.1126/science.adf1466">https://doi.org/10.1126/science.adf1466</a></li> </ul> <p>Please note the author order of this data repository is different from the author order and contributions in the review papers.</p> <p>REVIEW I: EVOLUTION, DISTRIBUTION AND USE</p> <ul> <li>catalogue_of_vascular_plants_of_madagascar.csv -- Comma-separated table with comprehensive taxonomic database of vascular plants of Madagascar, from the Catalogue of the Vascular Plants of Madagascar project. Contact: Peter Phillipson - peter.phillipson@mobot.org and Marina Rabarimanarivo - marina.rabarimanarivo@mobot.mg</li> <li>fungi_supplementary_material.zip -- Zipped archive containing: R script to get fungal endemism estimates from GBIF/UNITE and process data; comma-separated tables from GBIF, PlutoF, Goodman lichen checklist and Index Fungorum (as of 02/12/20): comma-separated table listing Madagascan taxa with endemism status. Contact: Rowena Hill - r.hill@kew.org</li> <li>lineages_madagascar.csv -- Comma-separated table with crown and stem ages, number of species, and geographic origin or distribution of sister clade of Malagasy endemic lineages, extracted from the literature. Contact: Jan Hackel - j.hackel@kew.org and Angelica Crottini - tiliquait@yahoo.it</li> <li>madagascar_fossil_genera_occurrences.csv -- Comma-separated table with worldwide fossil records for Malagasy species, downloaded from the PaleoBiology Database. -- Contact: Juan Carillo - juan.carrillo@mnhn.fr</li> <li>species_richness_modeling_occurrences.csv -- Comma-separated table with specimen-based occurrence data for Malagasy amphibians, grasses, lemurs, palms, reptiles, and Sarcolaenaceae. Contact: Weston Testo - westontesto@gmail.com</li> <li>taxon_description_by_year.csv -- Comma-separated table with years of basionym publication for Malagasy amphibians, reptiles, vascular plants, and ants. Contact: Weston Testo - westontesto@gmail.com</li> <li>vegetation_moatsmith_1km_extended.zip -- Raster geotiff with new expanded vegetation types based on Moat & Smith (2007). Contact: Justin Moat - j.moat@kew.org</li> <li>vertebrate_species_list.csv -- Comma-separated table with native species list and associated endemism of freshwater fishes, amphibians, reptiles, birds, and mammals of Madagascar (author-curated list based on data from The New Natural History of Madagascar and the IUCN Red List). Contact: Weston Testo - westontesto@gmail.com, Angelica Crottini - tiliquait@yahoo.it, Ferran Sayol - fsayol@gmail.com</li> </ul> <p>REVIEW II: THREATS AND OPPORTUNITIES</p> <ul> <li>catalogue_of_vascular_plants_of_madagascar.csv -- Comma-separated table with comprehensive taxonomic database of vascular plants of Madagascar, from the Catalogue of the Vascular Plants of Madagascar project. Contact: Peter Phillipson - peter.phillipson@mobot.org and Marina Rabarimanarivo - marina.rabarimanarivo@mobot.mg</li> <li>ex_situ_plants.csv -- Comma-separated table with numbers of ex situ conserved plant species, per family, from BGCI’s PlantSearch database and collections of Jardin Botanique Educatif and Parc Ivoloina. Contact: Malin Rivers - malin.rivers@bgci.org</li> <li>ex_situ_vertebrates.csv -- Excel spreadsheet with the list of extant native Malagasy vertebrates with information on their presence in at least one international zoo holding and whether they have been bred successfully over the last 12 months. Data from the Zoological Information Management (ZIM) Software performed in February 2021. Contact: Angelica Crottini - tiliquait@yahoo.it</li> <li>extinct_animals_madagascar.csv -- Comma-separated table of all known anthropogenic extinctions before 1500 CE in Madagascar. Contact: Ferran Sayol - fsayol@gmail.com</li> <li>madagascar_protected_areas_sources.csv -- Comma-separated table with comments and sources for columns in the protected area data in the csv and shapefile. Contact: Maria S. Vorontsova - m.vorontsova@kew.org</li> <li>madagascar_terrestrial_protected_areas.csv -- Comma-separated values with description of the pretected areas, matching the Protected Area Shapefile. This file contains French accents; correct display may depend on the software used. Contact: Daniel Edler - daniel.edler@umu.se and Henintsoa Razanajatovo - H.Razanajatovo@kew.org</li> <li>madagascar_terrestrial_protected_areas.zip -- ESRI Shapefile for the synthesized protected areas of Madagascar, including Key Biodiversity Areas and attributes. Contact: Daniel Edler - daniel.edler@umu.se and Rasolohery Andriambolantsoa - arasolohery@ileiry.com</li> <li>observed_and_predicted_threats.csv -- Comma-separated table with the number of species with each listed threat, as defined by the IUCN or predicted by our model, across taxonomic groups. Contact: Rob Cooke - 03rcooke@gmail.com</li> <li>phylogenetic_diversity_methods.zip -- Zipped archive containing community matrices, species range shapefiles, and R script used to estimate phylogenetic diversity for amphibians, mammals, and reptiles. Contact: Weston Testo - westontesto@gmail.com</li> <li>predicting_species_IUCN_status.zip -- Zipped archive containing data, scripts and an Rstudio project to: (1) prepare features for using IUCNN v1.0 to predict the conservation status for Not Evaluated species (01_feature_preparation); (2) predict species IUCN status assessment using neural networks (02_predicting_species_IUCN_status); (3) predict species’ threat status using neural networks (03_predicting_species_threats). Contact: Alexander Zizka - alexander.zizka@biologie.uni-marburg.de and Daniele Silvestro - daniele.silvestro@unifr.ch</li> <li>threat_predictions_iucnn.txt -- Tab-separated table with results of the conservation status prediction from a Bayesian Neural Network for 5,887 species of vascular plants from Madagascar. Values are the mean posterior probabilities for each IUCN Red List category. Contact: Daniele Silvestro - daniele.silvestro@unifr.ch</li> </ul>
Data Repository: 2022 Hawai'i Cesspool Hazard Assessment & Prioritization Tool
<p>Data Repository, Codebase, inputs and Results for the Hawaii Cesspool Prioritization Tool. A project conducted by University of Hawaii Sea Grant and Water Resources Research Center, Data updated October 2022. </p> <p>Please see also: <br> https://github.com/cshuler/Act132_Cesspool_Prioritization</p> <p>and </p> <p>https://health.hawaii.gov/wastewater/files/2022/11/prioritizationtoolreport.pdf</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.