Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
197
datasets available to search
ShareScore release 0.9.0
Dataset results
197 results for “Materials science”
Improving machine-learning models in materials science through large datasets
<p>1. Image of the <a href="https://alexandria.icams.rub.de/"><strong>Alexandria database </strong></a> state corresponding to the paper "<strong>Improving machine-learning models in materials science through large datasets</strong>".</p> <ul> <li>Static pbe calculations for 1D, 2D, 3D compounds can be found in 1D_pbe.tar.gz, 2D_pbe.tar.gz, 3D_pbe.tar.gz in batches of 100k materials. The latter also contains a separate convex hull pickle with all compounds on the pbe convex hull (convex_hull_pbe_2023.12.29.json.bz2) and a list of prototypes in the database (prototypes.json.bz2). The systematic 3D calculations performed for the article <strong>Improving machine-learning models in materials science through large datasets </strong>(in the paper referred to as round 2 and 3) can be found by the location keyword in the data dictionary of each ComputedStructureEntry containing "<strong>cgat_comp/quaternaries</strong>" (round 2) and "<strong>cgat_comp2/</strong>" (round 3). Round 1 (10.1002/adma.202210788) can be found under "cgat_comp/ternaries", ""cgat_comp/binaries".</li> <li>Static pbesol calculations for 3D compounds can be found in 3D_ps.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the pbesol convex hull (convex_hull_ps_2023.12.29.json.bz2). </li> <li>Static scan calculations for 3D compounds can be found in 3D_scan.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the scan convex hull (convex_hull_scan_2023.12.29.json.bz2). </li> <li>Geometry relaxation curves for 1D and 2D and 3D compounds calculated with PBE can be found in geo_opt_1D.tar.gz, geo_opt_2D.tar.gz. and geo_opt_3D.tar. Each file in each folder contains a batch of up to 10k relaxation trajectories.</li> <li>PBESOL relaxation trajectories for 3D compounds can be found in geo_opt_ps.tar</li> </ul> <p>2. Crystal graph attention networks to predict the volume (<a href="https://zenodo.org/api/records/12582650/draft/files/volume_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">volume_round_3.tar.gz</a>) and distance to the convex hull (<a href="https://zenodo.org/api/records/12582650/draft/files/e_above_hull_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">e_above_hull_round_3.tar.gz</a>) trained for the paper "Improving machine-learning models in materials science through large datasets".</p> <p>Can be used with the code at https://github.com/hyllios/CGAT/tree/main/CGAT.<br><strong>Note will predict the distance to the convex hull not normalized per atom when using the code on the github.<br></strong></p> <p>3. Alignn models as well as m3gnet and mace models corresponding to the publication can be found in <a href="https://zenodo.org/api/records/12582650/draft/files/alexandria_v2.tar.gz/content" target="_blank" rel="noopener noreferrer">alexandria_v2.tar.gz</a></p> <p>4. scripts.tar.gz Some scripts used for generating CGAT input data/ performing parallel predictions and for relaxations with m3gnet/mace force fields</p>
Si data files for Galaxy materials science tutorials
<p>This is a training dataset for use in Galaxy materials science tutorials. These files can be used to demonstrate the AIRSS (Ab-Initio Random Structure Searching) method for finding muon stopping sites, using the UEP (Unperturbed Electrostatic Potential) technique for the optimisation stage of that method.</p> <p>The files included are:</p> <ul> <li><strong>Si.cell:</strong> structure file containing atom locations</li> <li><strong>Si.den_fmt:</strong> electron density data, generated with CASTEP</li> <li><strong>Si.castep:</strong> CASTEP log file for the electron density calculation</li> <li><strong>Si-muairss-uep.yaml:</strong> configuration file for the AIRSS / UEP workflow</li> </ul>
Course Materials for Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730)
In today's world, understanding environmental data and making informed decisions based on it is crucial for addressing complex environmental challenges. Yale School of the Environment's Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730) course serves as an introduction to the integration of environmental data using R programming language, coupled with machine learning techniques. This dataset contains a zip file with all the data files used in this course, along with a README that has the metadata for those files.
Survey on Open Science Practices in Functional Neuroimaging. Dataset and Materials
<p>Preregistration of the hypotheses and methods of an empirical study before analysis, the sharing of primary research data, and compliance with data standards such as the Brain Imaging Data Structure (BIDS), are considered effective practices to secure progress and to substantiate quality of research. We investigated the current level of adoption of open science practices in neuroimaging and the difficulties that prevent researchers from using them. A PubMed search with the search terms ("fMRI" OR "functional magnetic resonance imaging" OR "functional Magnetic Resonance Imaging") was done to collect email addresses from corresponding authors of scientific articles published between 2010/01/01 and 2020/08/28. An email was sent to 14,690 addresses on 2020/01/12 with an invitation to participate, including a personalized link to the survey. If the recipients did not click the link or did not complete the survey after 14 days, they received a single reminder email. The questionnaire was composed of five building blocks. The Blocks 1-3 focused on three areas of open science practices: data structure, preregistration and data sharing. The fourth block asked about technical expertise with software and the fifth part assessed sociodemographic data.</p>
Open Science and Authorship of Supplementary Material for the MES research community
<p>This spreadsheet contains the data and the results from the analysis described in the paper "Open Science and Authorship of Supplementary Material. Evidence from a Research Community." being accepted at STI 2022.</p>
Reproducible and Attributable Materials Science Workflows
<p>This set includes the deidentified data, reproducible analysis and research report of the project on Reproducible and Attributable Materials Science Workflows.</p>
Supporting Material for article "The ELIXIR Core Data Resources: fundamental infrastructure for the life sciences"
<p>This data set is the Supporting Material referred to in the Supplementary Data for the article "The ELIXIR Core Data Resources: fundamental infrastructure for the life sciences" (Drysdale, et al.) submitted for publication in April 2019.</p> <p> </p>
Evaluating Open Science Practices in Indoor Positioning and Indoor Navigation Research (Supplementary Material: Full Paper Listing and Analysis)
<p>Supplementary material of the paper:</p> <p>Title: "Evaluating Open Science Practices in Indoor Positioning and Indoor Navigation Research"<br>Subtitle: "A Survey of the IPIN's Reference Papers of 2022 and 2023 Editions"</p> <p>The paper is accepted to the "14th International Conference on Indoor Positioning and Indoor Navigation, IPIN 2024, Hong Kong, October 14-17, 2024, IEEE, 2024.</p> <p>An Author's accepted version of the manuscript is available here: <a href="../records/13684170" target="_blank" rel="noopener">https://zenodo.org/records/13684170</a> </p> <p>If you want to refer to this work, please cite this Zenodo entry as well as the published conference version.</p> <p> </p> <p>---------------------------------------</p> <p>This entry contains two files:</p> <ul> <li>"Paper Characterization Spreadsheet.xlsx": <strong>The spreadsheet of the full analysis of this work</strong>, as described in the paper. It characterizes various features of the analyzed papers and forms the raw data on which the analyses of our work were based.</li> <li>"Main features of the manuscripts analysed in Zenodo Record #12088175.pdf": A document summarizing the main features of the IPIN's Reference Papers of the 2022 and 2023 Editions, that contain some form of open resources (Open Data, Code, or Material).</li> </ul> <p> </p> <p> </p> <p> </p>
Seeing Through the "Science Eyes" of the ExoMars Rover - Supplementary Material
<p>Simulated views from ExoMars PanCam instrument to assist operations planning.</p> <p>Described in more detail in the linked journal article.</p>
MuSpinSim data files for Galaxy materials science tutorials
<p>This is a training dataset for use in Galaxy materials science tutorials. These files can be compared to the output of simulations by MuSpinSim for dissipation of muon spins.</p> <p>The files included are:</p> <ul> <li><strong>dissipation_theory.dat:</strong> theoretical values formatted as a MuSpinSim output</li> <li><strong>experiment.dat:</strong> mock experimental values formatted as a MuSpinSim output</li> </ul>
Materials Science Optimization Benchmark Dataset for Multi-Objective, Multi-Fidelity Optimization of Hard-Sphere Packing Simulations
<p>Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a “Turing test” of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In the fields of materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 494498 random hard-sphere packing simulations representing 206 CPU days worth of computational overhead were performed across nine input parameters with linear constraints and two discrete fidelities each with continuous fidelity parameters and results were logged to a free-tier shared MongoDB Atlas database. Two core tabular datasets resulted from this study: 1. a failure probability dataset containing unique input parameter sets and the estimated probabilities that the simulation will fail at each of the two steps, and 2. a regression dataset mapping input parameter sets (including repeats) to particle packing fractions and computational runtimes for each of the two steps. These two datasets are used to create a surrogate model as close as possible to running the actual simulations by incorporating simulation failure and heteroskedastic noise. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This is in contrast with a more traditional approach that imposes a-priori assumptions such as Gaussian noise e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.</p> <p>For usage instructions, see https://matsci-opt-benchmarks.readthedocs.io/.</p>
Supplementary material 10: Institutional collection dashboard: specimens from the collection of the California Academy of Sciences (CAS) from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063
Dashboard charts showing only specimens from the collection of the California Academy of Sciences. This page shows data from species-rank treatments. When viewed using a browser (such as Google Chrome) with an internet connection, this page sends a series of queries to Plazi and integrates the results with the Google Charts API to produce 37 interactive dashboard charts.
Supplementary material for the publication: "Efficient Surrogate Models for Materials Science Simulations: Machine Learning-based Prediction of Microstructure Properties"
<p><span><span><span>This dataset contains supplementary code, images and models for the publication „Efficient Surrogate Models for Materials Science Simulations: Machine Learning-based Prediction of Microstructure Properties“.</span></span></span></p> <p> </p> <p><span><span><span>The content will be updated and additionally linked to the corresponding git repositories.</span></span></span></p>
HoneyBee: Progressive Instruction Finetuning of Large Language Models for Materials Science
<p>We propose an instruction-based process for trustworthy data curation in materials science (MatSci-Instruct), which we then apply to finetune a LLaMa-based language model targeted for materials science (HoneyBee). MatSci-Instruct helps alleviate the scarcity of relevant, high-quality materials science textual data available in the open literature, and HoneyBee is the first billion-parameter language model specialized to materials science. In MatSci-Instruct we improve the trustworthiness of generated data by prompting multiple commercially available large language models for generation with an Instructor module (e.g. Chat-GPT) and verification from an independent Verifier module (e.g. Claude). Using MatSci-Instruct, we construct a dataset of multiple tasks and measure the quality of our dataset along multiple dimensions, including accuracy against known facts, relevance to materials science, as well as completeness and reasonableness of the data. Moreover, we iteratively generate more targeted instructions and instruction-data in a finetuning-evaluation-feedback loop leading to progressively better performance for our finetuned HoneyBee models. Our evaluation on the MatSci-NLP benchmark shows HoneyBee's outperformance of existing language models on materials science tasks and iterative improvement in successive stages of instruction-data refinement. We study the quality of HoneyBee's language modeling through automatic evaluation and analyze case studies to further understand the model's capabilities and limitations. Our code and relevant datasets are publicly available at https://github.com/BangLab-UdeM-Mila/NLP4MatSci-HoneyBee.</p>
Dataset - A Python-Based Approach to Sputter Deposition Simulations in Combinatorial Materials Science
<p>This dataset accompanies the publication <em>"A Python-Based Approach to Sputter Deposition Simulations in Combinatorial Materials Science,"</em> which presents and validates pySIMTRA, a Python wrapper for the Monte Carlo-based SIMTRA simulation tool. The dataset includes all measured and simulated data shown in the publication, as well as additional animations visualizing the compositions in the multinary composition space.</p> <p>The dataset contains the compositions for each of the seven materials libraries (in at.%) in the quaternary Ni-Pd-Pt-Ru system. Additionally, it provides the simulated number of particles as outputted by SIMTRA, which serve as the basis for composition estimation. Both the compositional data and the particle counts are supplied in .csv format. To supplement the results, 3D animations of the quaternary compositional spaces are included, showing the comparison between simulated compositions (red dots) and measured compositions (blue dots). These animations offer a more intuitive visualization of the data compared to the static Figures in the publication and are supplied as .gif files.</p> <p>Due to the in-depth analysis of cathode tilt discussed in the paper, the dataset also includes simulation results for the ternary Pd-Pt-Ru library, highlighting the effect of varying the cathode tilt angle. Simulations were conducted for tilt angles of 10°, 9.5°, 9°, and 8.5°.</p>
Supporting material for "Impact of gender on the formation and outcome of formal mentoring relationships in the life sciences"
<p>This repository contains data and analysis code associated with the manuscript: L.P. Schwartz, J. Liénard, S. V. David. (2022) "Impact of gender on formation and outcome of formal mentoring relationships in the life sciences." Figures and tables in the manuscript can be produced by running the make_figures.ipynb notebook. Figures have been marked with headings indicating their position in the manuscript (Figure 1, Figure S1, etc.). In addition, the notebook contains code to reproduce regression analyses that are cited in the text but not directly associated with a figure.</p> <p>Data on mentoring relationships derives from Academic Family Tree (AFT, www.academictree.org) and public data sources on funding, publications, and awards. Inclusion criteria, public data sources, and procedures for linking across sources are described in the manuscript. Personal identifiers for researchers have been anonymized, but remain consistent across all data in the repository. In other words, the personal identifier "1" refers to the same person in all dataframes in the repository. But, that person is *not* the same researcher identified as "1" on the public AFT website.</p> <p><strong>Installation</strong></p> <p>Requires Python 3.x. and Pandas. To load required libraries using Anaconda, run:</p> <p>`conda create --name aft -c conda-forge pandas numpy scipy ipython jupyterlab scipy scikit-learn pandas matplotlib numpy statsmodels seaborn pytables`</p> <p><strong>Dataframes</strong></p> <p>Data is stored as a series of Pandas dataframes within HDF5 or CSV files:</p> <p>* cng_tc: The primary dataset used in the analysis. The name is an acronym for "connections" (i.e. training relationships, "cn"), "gender" ("g"), and "trainee count" ("tc"). Each row contains data on the mentor and trainee in one training relationship. See manuscript for inclusion criteria.</p> <p>* mentors: Data on mentors. Each row contains data on one mentor. See manunscript for inclusion criteria.</p> <p>* mentors_grants, mentors_hindex, mentors_locs_ranked: Subset of mentors with data available for funding (mentors_grants), citation (mentors_hindex), and institution rank (mentors_locs_ranked).</p> <p>* mentors_nobel, mentors_hhmi, mentors_nas: Subsets of mentors that received a Nobel (mentors_nobel), Howard Hughes Medical Institute grants (mentors_hhmi), or membership in the National Academy of Sciences (mentors_nas). See manuscript for details of data sources and linking procedures.</p> <p>* cn, cng, first_names, gn, gn_all, locs: Partial data (connections only, inferred gender only, connections and gender only, location only, first names and inferred gender only) for more inclusive sets of researchers in AFT. They are generally not used used for analysis, but have been included here to calculate statistics on the total amount of data included and to screen for data from U.S. locations.</p> <p>* nsf_gender_phds, nsf_gender_pds: National Science Foundation survey data on gender and fraction PhDs conferred per year (nsf_gender_phds) or fraction postdocs employed per year (nsf_gender_pds). See manuscript for details of data source.</p> <p>* photo: Data for validation of gender inference method.</p> <p><strong>Dataframe columns</strong></p> <p>* amount: Mentor's total funding<br> * amount_adj: Mentor's total funding (adjusted to 2020 dollars)<br> * broad_field: Mentor's general research area (e.g., life sciences, engineering, based on National Science Foundation classifications)<br> * continue: Whether trainee went on to become a mentor (i.e., has trainees listed in AFT)<br> * country: Country in which mentor's current institution is located<br> * firstname: First name of researcher (table of first names is not aligned with tables containing anonymized personal identifiers)<br> * first_grant_year: Year of mentor's first grant<br> * funding_rate: Mentor's annual funding rate (since first grant)<br> * funding_rate_adj: Mentor's annual funding rate (since first grant) adjusted to 2020 dollars<br> * hhmi: Whether mentor was granted HHMI funding<br> * hindex: Mentor's hindex<br> * location: Name of mentor's current institution<br> * locid: Identifier for mentor's institution<br> * locid_rank: Postion of mentor's institution in 2015 Quacquarelli-Symonds rankings (lower numbers are better)<br> * locid_rank_rev: Reversed version of "locid_rank" (i.e., higher numbers are better)<br> * majorarea: Mentor's specific research area (e.g, neuroscience)<br> * male_mentor, male trainee: Whether the probability that a researcher's first name is used by a person identifying as a man meets threshold (see manuscript for details on gender inference using first names)<br> * match_score: Score for string match between institution or name of awardee and researcher<br> * mentor_career_start: The date at which the mentor's academic career began<br> * mentor_continue_rate: Fraction of mentor's trainees that become mentors<br> * mentor_continue_rate_ft: Fraction of mentor's woman trainees that become mentors<br> * mentor_continue_rate_mt: Fraction of mentor's man trainees that become mentors<br> * mentor_t_p_male0: Fraction of mentor's trainees that are men<br> * mentor_t_p_male0_gs: Fraction of mentor's trainees that are men (graduate students only)<br> * mentor_t_p_male0_pd: Fraction of mentor's trainees that are men (postdocs only)<br> * mentor_tcount0: Mentor's total number of trainees<br> * nas: Whether mentor is a member of the National Academy of Sciences<br> * nobel: Whether mentor is a Nobel laureate<br> * p_male_mentor, p_male_trainee: Probability that a researcher's first name is used by a person identifying as a man<br> * pid: Anonymized identifier of researcher<br> * pid_mentor: Anonymized identifier of mentor in training relationship<br> * pid_trainee: Anonymized identifier of trainee in training relationship<br> * pq: "1" if data on training relationship is drawn from ProQuest database and has not been manually edited a human AFT user<br> * relation: Type of training relationship (1: graduate student, 2: postdoc)<br> * scorer1, scorer2, scorer3: Results of photo validation of gender inference for each scorer<br> * start: Training start year<br> * stop: Training end year<br> * trainee_tcount: Total people that the trainee has trained<br> * triad: Whether trainee has participated in both a graduate-level and postdoctoral training relationship</p> <p>The cn dataframe follows slightly different naming conventions, but is not generally used in the analysis (pid1 = pid_trainee, pid2 = pid_mentor, startdate = start, stopdate = stop).</p>
Using ELN Functionality of Kadi4Mat (KadiWeb) in a Materials Science Case Study of a User Facility
<p>This record contains the dataset belonging to the paper "Using ELN Functionality of Kadi4Mat (KadiWeb) in a Materials Science Case Study of a User Facility"</p>
Yes! We're open. Open science and the future of academic practices in translation and interpreting studies - Supplementary material
<p>Supplementary material to the article "<em>Yes! We’re open</em>. Open science and the future of academic practices in translation and interpreting studies" by Christian Olalla-Soler. </p> <ul> <li>Sheet 1: Translation and Interpreting Studies journals and bibliometric indicators.</li> <li>Sheet 2: Translation and Interpreting Studies articles in Scopus.</li> <li>Sheet 3: Pre-registrations related to translation and interpreting.</li> </ul> <p>Reference:</p> <p>Olalla-Soler, Christian (2021). "<em>Yes! We’re open</em>. Open science and the future of academic practices in translation and interpreting studies". <em>Translation & Interpreting</em> 13 (2): 1-28. <a href="https://doi.org/10.12807/ti.113202.2021.a01">https://doi.org/10.12807/ti.113202.2021.a01</a></p>
Figs 59‒62 in Revision of type and non-type material assigned to the genus Orthocladius by Goetghebuer (1940-1950), deposited in the Royal Belgian Institute of Natural Sciences (Diptera: Chironomidae)
Figs 59‒62. Orthocladius (Pogonocladius) consobrinus (Holmgren, 1869) (= Orthocladius crassicornis Goetghebuer, 1937). 59 ‒ head, 60 ‒ inferior volsella, 61 ‒ inferior volsella of lectotype of P. consobrinus from NHRM; 62 ‒ gonostylus.
Figs 47‒52 in Revision of type and non-type material assigned to the genus Orthocladius by Goetghebuer (1940-1950), deposited in the Royal Belgian Institute of Natural Sciences (Diptera: Chironomidae)
Figs 47‒52. Hydrobaenus corax (Kieffer, 1924). 47 ‒ pulvilli, 48 ‒ hypopygium, 49 ‒ anal point, 50 ‒ virga, 51 ‒ inferior volsella, 52 ‒ gonostylus.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.