Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,025
datasets available to search
ShareScore release 0.9.0
Dataset results
6,025 results for “science”
Improving machine-learning models in materials science through large datasets
<p>1. Image of the <a href="https://alexandria.icams.rub.de/"><strong>Alexandria database </strong></a> state corresponding to the paper "<strong>Improving machine-learning models in materials science through large datasets</strong>".</p> <ul> <li>Static pbe calculations for 1D, 2D, 3D compounds can be found in 1D_pbe.tar.gz, 2D_pbe.tar.gz, 3D_pbe.tar.gz in batches of 100k materials. The latter also contains a separate convex hull pickle with all compounds on the pbe convex hull (convex_hull_pbe_2023.12.29.json.bz2) and a list of prototypes in the database (prototypes.json.bz2). The systematic 3D calculations performed for the article <strong>Improving machine-learning models in materials science through large datasets </strong>(in the paper referred to as round 2 and 3) can be found by the location keyword in the data dictionary of each ComputedStructureEntry containing "<strong>cgat_comp/quaternaries</strong>" (round 2) and "<strong>cgat_comp2/</strong>" (round 3). Round 1 (10.1002/adma.202210788) can be found under "cgat_comp/ternaries", ""cgat_comp/binaries".</li> <li>Static pbesol calculations for 3D compounds can be found in 3D_ps.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the pbesol convex hull (convex_hull_ps_2023.12.29.json.bz2). </li> <li>Static scan calculations for 3D compounds can be found in 3D_scan.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the scan convex hull (convex_hull_scan_2023.12.29.json.bz2). </li> <li>Geometry relaxation curves for 1D and 2D and 3D compounds calculated with PBE can be found in geo_opt_1D.tar.gz, geo_opt_2D.tar.gz. and geo_opt_3D.tar. Each file in each folder contains a batch of up to 10k relaxation trajectories.</li> <li>PBESOL relaxation trajectories for 3D compounds can be found in geo_opt_ps.tar</li> </ul> <p>2. Crystal graph attention networks to predict the volume (<a href="https://zenodo.org/api/records/12582650/draft/files/volume_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">volume_round_3.tar.gz</a>) and distance to the convex hull (<a href="https://zenodo.org/api/records/12582650/draft/files/e_above_hull_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">e_above_hull_round_3.tar.gz</a>) trained for the paper "Improving machine-learning models in materials science through large datasets".</p> <p>Can be used with the code at https://github.com/hyllios/CGAT/tree/main/CGAT.<br><strong>Note will predict the distance to the convex hull not normalized per atom when using the code on the github.<br></strong></p> <p>3. Alignn models as well as m3gnet and mace models corresponding to the publication can be found in <a href="https://zenodo.org/api/records/12582650/draft/files/alexandria_v2.tar.gz/content" target="_blank" rel="noopener noreferrer">alexandria_v2.tar.gz</a></p> <p>4. scripts.tar.gz Some scripts used for generating CGAT input data/ performing parallel predictions and for relaxations with m3gnet/mace force fields</p>
Datasets from Analysis of the strategic management of science, research and innovation thesis
<p>Datasets were obtained from Czech R&D information system and used in my diploma thesis. The thesis (in Czech language) explores system of governance in R&D in the Czech republic and his impacts on the field of molecular biology in the period of 1995-2014.</p>
Dataset: Rainbow color map distorts and misleads research in hydrology – guidance for better visualizations and science communication
<p>The rainbow color map is scientifically incorrect and hinders people with color vision deficiency to view visualizations in a correct way. Due to perceptual non-uniform color gradients within the rainbow color map the data representation is distorted what can lead to misinterpretation of results and flaws in science communication. Here we present the data of a paper survey of 797 scientific publication in the journal Hydrology and Earth System Sciences. With in the survey all papers were classified according to color issues. Find details about the data below.</p> <ul> <li><code>year</code> = year of publication (YYYY)</li> <li><code>date</code> = date (YYYY-MM-DD) of publication</li> <li><code>title</code> = full paper title from journal website</li> <li><code>authors</code> = list of authors comma-separated</li> <li><code>n_authors</code> = number of authors (integer between 1 and 27)</li> <li><code>col_code</code> = color-issue classification (see below)</li> <li><code>volume</code> = Journal volume</li> <li><code>start_page</code> = first page of paper (consecutive)</li> <li><code>end_page</code> = last page of paper (consecutive)</li> <li><code>base_url</code> = base url to access the PDF of the paper with <code>/volume/start_page/year/</code></li> <li><code>filename</code> = specific file name of the paper PDF (e.g. <code>hess-9-111-2005.pdf</code>)</li> </ul> <p>Color classification is stored in the <code>col_code</code> variable with:</p> <ul> <li><code>0</code> = chromatic and issue-free,</li> <li><code>1</code> = red-green issues,</li> <li><code>2</code>= rainbow issues and</li> <li><code>bw</code>= black and white paper.</li> </ul> <p> </p> <p>See more details (e.g., sample code to analyse the survey data) on https://github.com/modche/rainbow_hydrology</p> <p>Paper: Stoelzle, M. and Stein, L.: Rainbow color map distorts and misleads research in hydrology – guidance for better visualizations and science communication, Hydrol. Earth Syst. Sci., 25, 4549–4565, https://doi.org/10.5194/hess-25-4549-2021, 2021.</p> <p> </p> <p> </p>
LamaH-CE: LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe – files
<p><strong>Version 1.0 - This version is the final revised one.</strong></p> <p>This is the LamaH-CE dataset accompanying the paper: Klingler et al., LamaH-CE | LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe, published at Earth System Science Data (ESSD), 2021 (<a href="https://doi.org/10.5194/essd-13-4529-2021">https://doi.org/10.5194/essd-13-4529-2021</a>).</p> <p>LamaH-CE contains a collection of runoff and meteorological time series as well as various (catchment) attributes for 859 gauged basins. The hydrometeorological time series are provided with daily and hourly time resolution including quality flags. All meteorological and the majority of runoff time series cover a span of over 35 years, which enables long-term analyses with high temporal resolution.<br> LamaH is in its basics quite sililar to the well-known CAMELS datasets for the contiguous United States (<a href="https://doi.org/10.5194/hess-21-5293-2017">https://doi.org/10.5194/hess-21-5293-2017</a>), Chile (<a href="https://doi.org/10.5194/hess-22-5817-2018">https://doi.org/10.5194/hess-22-5817-2018</a>), Brazil (<a href="https://doi.org/10.5194/essd-12-2075-2020">https://doi.org/10.5194/essd-12-2075-2020</a>), Great Britain (<a href="https://doi.org/10.5194/essd-12-2459-2020">https://doi.org/10.5194/essd-12-2459-2020</a>) and Australia (<a href="https://doi.org/10.5194/essd-13-3847-2021">https://doi.org/10.5194/essd-13-3847-2021</a>), but new features like additional basin delineations (intermediate catchments) and attributes allow to consider the hydrological network and river topology in further applications.</p> <p>We provide two different files to download: 1) Hydrometeorological time series with daily and hourly resolution, which requires decompressed about 70 GB of free disk space. 2) Hydrometeorological time series only with daily resolution, which requires 5 GB. Beyond the temporal resolution of the time series, there are no differences.</p> <p><strong>Note: </strong>It is recommended to read the supplementary info file before using the dataset. For example, it clarifies the time conventions and that <strong>NAs</strong> are indicated by the number<strong> -999</strong> in the <strong>runoff time series</strong>.</p> <p><strong>Disclaimer:</strong> We have created LamaH with care and checked the outputs for plausibility. By downloading the dataset, you agree that we nor the provider of the used source datasets (e.g. runoff time series) cannot be liable for the data provided. The runoff time series of the German federal states Bavaria and Baden-Württemberg are retrospective checked and updated by the hydrographic services. Therefore, it might be appropriate to obtain more up-to-date runoff data from Bavaria (<a href="https://www.gkd.bayern.de/en/rivers/discharge/tables">https://www.gkd.bayern.de/en/rivers/discharge/tables</a>) and Baden-Württemberg (<a href="https://udo.lubw.baden-wuerttemberg.de/public/p/pegel_messwerte_leer">https://udo.lubw.baden-wuerttemberg.de/public/p/pegel_messwerte_leer</a>). Runoff data from the Czech Republic may not be used to set up operational warning systems (<a href="https://www.chmi.cz/files/portal/docs/hydro/denni_data/Podminky_uziti.pdf">https://www.chmi.cz/files/portal/docs/hydro/denni_data/Podminky_uziti.pdf</a>).</p> <p><strong>License: </strong>This work is licensed with CC BY-SA 4.0 (<a href="https://creativecommons.org/licenses/by-sa/4.0/">https://creativecommons.org/licenses/by-sa/4.0/</a>). This means that you may freely use and modify the data (even for commercial purposes). But you have to give appropriate credit (associated ESSD paper, version of dataset and all sources which are declared in the folder "Info"), indicate if and what changes were made and distribute your work under the same public license as the original.</p> <p><strong>Additional references: </strong>We ask kindly for compliance in citing the following references when using LamaH, as an agreement to cite was usually a condition of sharing the data: BAFU (2020), CHMI (2020), GKD (2020), HZB (2020), LUBW (2020), BMLFUW (2013), Broxton et al. (2014), CORINE (2012), EEA (2019), ESDB (2004), Farr et al. (2007), Friedl and Sulla-Menashe (2019), Gleeson et al. (2014), HAO (2007), Hartmann and Moosdorf (2012), Hiederer (2013a, b), Linke et al. (2019), Muñoz Sabater et al. (2021), Muñoz Sabater (2019a), Myneni et al. (2015), Pelletier et al. (2016), Toth et al. (2017), Trabucco and Zomer (2019), and Vermote (2015). These references are listed in detail in the accompanying <a href="https://doi.org/10.5194/essd-13-4529-2021">paper</a>.</p> <p><strong>Supplements: </strong>We have created additional files after publication (therefore non peer-reviewed):<br> 1) Shapefiles for reservoirs (points) and cross-basin water transfers (lines) including several attributes as well as tables with information about the accumulated storage volume and effective catchment area (considerung artificial in- and outflows) for every runoff gauge.<br> 2) Water quality data (e.g. dissolved oxygen, water temperature, conductivity, NO3-N), which are suitable to the gauges. The data for water quality may not be used for commercial purposes.<br> If you are interessted, just send us an email with your name, affiliation and the intended purpose for the requested files to the address listed below. If you find any errors in the dataset, feel free to send us an email to: christoph.klingler@boku.ac.at</p>
RDA Overview for the Social Sciences - October 2022
<p>This upload is the fourth version of a dataset on the RDA groups and outputs which was generated in October 2022. It is connected to a report that described a project that assessed the RDA Groups and Outcomes and their relevance for social science researchers. This datasets gives an overview of the results of this analysis, providing a list of all RDA Groups as well as an indication of the relevance for Social Sciences researchers. The project was executed and updated by Data Archiving and Networked Services (DANS).</p> <p>This dataset contains:</p> <p>- an excel overview of the RDA Working and Interest groups</p> <p>- additional csv copies of the separate sheets of the excel overview</p>
Products and Models for "Early Release Science of the Exoplanet WASP-39b with JWST NIRCam"
<p>Associated Publication: <a href="https://www.nature.com/articles/s41586-022-05590-4">https://www.nature.com/articles/s41586-022-05590-4</a><br> <br> OVERVIEW: Measuring the metallicity and carbon-to-oxygen (C/O) ratio in exoplanet atmospheres is a fundamental step towards constraining the dominant chemical processes at work and, if in equilibrium, revealing planet formation histories. Transmission spectroscopy<sup> </sup>provides the necessary means by constraining the abundances of oxygen- and carbon-bearing species; however, this requires broad wavelength coverage, moderate spectral resolution, and high precision that, together, are not achievable with previous observatories. Now that JWST has commenced science operations, we are able to observe exoplanets at previously uncharted wavelengths and spectral resolutions. Here we report time-series observations of the transiting exoplanet WASP-39b using JWST’s Near InfraRed Camera (NIRCam). The long-wavelength spectroscopic and short-wavelength photometric light curves span 2.0 – 4.0 µm, exhibit minimal systematics, and reveal well-defined molecular absorption features in the planet’s spectrum. Specifically, we detect gaseous H<sub>2</sub>O in the atmosphere and place an upper limit on the abundance of CH<sub>4</sub>. The otherwise prominent CO<sub>2</sub> feature at 2.8 µm is largely masked by H<sub>2</sub>O. The best-fit chemical equilibrium models favour an atmospheric metallicity of 1–100× solar (i.e., an enrichment of elements heavier than helium relative to the Sun) and a sub-stellar carbon-to-oxygen (C/O) ratio. The inferred high metallicity and low C/O ratio may indicate significant accretion of solid materials during planet formation<sup> </sup>or disequilibrium processes in the upper atmosphere.</p>
[Dataset] Does Volunteer Engagement Pay Off? An Analysis of User Participation in Online Citizen Science Projects
<p><strong>Explanation/Overview:</strong></p> <p>Corresponding dataset for the analyses and results achieved in the CS Track project in the research line on participation analyses, which is also reported in the publication "Does Volunteer Engagement Pay Off? An Analysis of User Participation in Online Citizen Science Projects", a conference paper for the conference CollabTech 2022: <a href="https://link.springer.com/book/10.1007/978-3-031-20218-6">Collaboration Technologies and Social Computing</a> and published as part of the <a href="https://link.springer.com/bookseries/558">Lecture Notes in Computer Science</a> book series (LNCS,volume 13632) <a href="https://link.springer.com/chapter/10.1007/978-3-031-20218-6_5">here</a>. The usernames have been anonymised.</p> <p><strong>Purpose:</strong></p> <p>The purpose of this dataset is to provide the basis to reproduce the results reported in the associated deliverable, and in the above-mentioned publication. As such, it <strong>does not</strong> represent <strong>raw data</strong>, but rather files that already include certain analysis steps (like calculated degrees or other SNA-related measures), ready for analysis, visualisation and interpretation with R.</p> <p><strong>Relatedness:</strong></p> <p>The data of the different projects was derived from the forums of 7 Zooniverse projects based on similar discussion board features. The projects are: 'Galaxy Zoo', 'Gravity Spy', 'Seabirdwatch', 'Snapshot Wisconsin', 'Wildwatch Kenya', 'Galaxy Nurseries', 'Penguin Watch'.</p> <p><strong>Content:</strong></p> <p>In this Zenodo entry, several files can be found. The structure is as follows (<code>files</code> and <strong>folders </strong>and<strong> </strong><em>descriptions</em>).</p> <ul> <li><code>corresponding_calculations.html</code> <ul> <li><em>Quarto-notebook to view in browser</em></li> </ul> </li> <li><code>corresponding_calculations.qmd</code> <ul> <li><em>Quarto-notebook to view in RStudio</em></li> </ul> </li> <li><strong>assets</strong> <ul> <li><strong>data</strong> <ul> <li><strong>annotations</strong> <ul> <li><code>annotations.csv</code> <ul> <li><em>List of annotations made per day for each of the analysed projects</em></li> </ul> </li> </ul> </li> <li><strong>comments</strong> <ul> <li><code>comments.csv </code> <ul> <li><em>Total list of comments with several data fields (i.e., comment id, text, reply_user_id)</em></li> </ul> </li> </ul> </li> <li><strong>rolechanges</strong> <ul> <li><code>478_rolechanges.csv</code> <ul> <li><em>List of roles per user to determine number of role changes </em></li> </ul> </li> <li><code>1104_rolechanges.csv</code> <ul> <li><em>...</em></li> </ul> </li> <li><code>...</code></li> </ul> </li> <li><strong>totalnetworkdata</strong> <ul> <li><strong>Edges</strong> <ul> <li><code>478_edges.csv</code> <ul> <li><em>Network data (edge set) for the given projects (without time slices)</em></li> </ul> </li> <li><code>1104_edges.csv</code> <ul> <li><em>...</em></li> </ul> </li> <li><code>...</code></li> </ul> </li> <li><strong>Nodes</strong> <ul> <li><code>478_nodes.csv</code> <ul> <li><em>Network data (node set) for the given projects (without time slices)</em></li> </ul> </li> <li><code>1104_nodes.csv</code> <ul> <li><em>...</em></li> </ul> </li> <li><code>...</code></li> </ul> </li> </ul> </li> <li><strong>trajectories</strong> <ul> <li><em>Network data (edge and node sets) for the given projects and all time slices (Q1 2016 - Q4 2021)</em></li> <li><strong>478</strong> <ul> <li><strong>Edges</strong> <ul> <li> <p><code>edges_4782016_q1.csv</code></p> </li> <li> <p><code>edges_4782016_q2.csv</code></p> </li> <li> <p><code>edges_4782016_q3.csv</code></p> </li> <li> <p><code>edges_4782016_q4.csv</code></p> </li> <li> <p><code>...</code></p> </li> </ul> </li> <li><strong>Nodes</strong> <ul> <li><code>nodes_4782016_q1.csv</code></li> <li> <p><code>nodes_4782016_q4.csv</code></p> </li> <li> <p><code>nodes_4782016_q3.csv</code></p> </li> <li> <p><code>nodes_4782016_q2.csv</code></p> </li> <li> <p><code>...</code></p> </li> </ul> </li> </ul> </li> <li> <p><strong>1104</strong> </p> <ul> <li> <p><strong>Edges</strong> </p> <ul> <li> <p><code>...</code></p> </li> </ul> </li> <li> <p><strong>Nodes</strong> </p> <ul> <li> <p><code>...</code></p> </li> </ul> </li> </ul> </li> <li> <p>...</p> </li> </ul> </li> </ul> </li> <li><strong>scripts</strong> <ul> <li><code>datavizfuncs.R</code> <ul> <li><em>script for the data visualisation functions, automatically executed from within </em><code>corresponding_calculations.qmd</code></li> </ul> </li> <li><code>import.R</code> <ul> <li><em>script for the import of data, automatically executed from within </em><code>corresponding_calculations.qmd</code></li> </ul> </li> </ul> </li> </ul> </li> <li><strong>corresponding_calculations_files</strong> <ul> <li>f<em>iles for the html/qmd view in the browser/RStudio</em></li> </ul> </li> </ul> <p><strong>Grouping:</strong></p> <p>The data is grouped according to given criteria (e.g., <code>project_title </code>or <code>time</code>). Accordingly, the respective files can be found in the data structure</p>
Si data files for Galaxy materials science tutorials
<p>This is a training dataset for use in Galaxy materials science tutorials. These files can be used to demonstrate the AIRSS (Ab-Initio Random Structure Searching) method for finding muon stopping sites, using the UEP (Unperturbed Electrostatic Potential) technique for the optimisation stage of that method.</p> <p>The files included are:</p> <ul> <li><strong>Si.cell:</strong> structure file containing atom locations</li> <li><strong>Si.den_fmt:</strong> electron density data, generated with CASTEP</li> <li><strong>Si.castep:</strong> CASTEP log file for the electron density calculation</li> <li><strong>Si-muairss-uep.yaml:</strong> configuration file for the AIRSS / UEP workflow</li> </ul>
Hyperspectral Imaging dataset for use in Heritage Science
<p>The following data sets were collected to support the potential uses of opensource data in the context of digital humanities and heritage sciences. </p> <p>This proposed experiment is conducted by the UCL Institute for Sustainable Heritage in collaboration with the Centre for Digital Humanities. Imaging methods including Photography, Multispectral Imaging, Hyperspectral Imaging and Xray Fluorescence Mapping have been collected along with the complete readout metadata of the instrumentation. </p> <p>We hope that you find the data helpful, and we welcome you to use the data in any way you wish, for all and any analysis development purposes. For us to build upon this research, we ask that in return you would be willing to share in some regard your experiences in using open-source data, using our data, successes and issues. </p> <p>If you would be willing to engage with us in this endeavor, please feel free to contact us so that we may be able to follow up with you. </p> <p>Other Data sets available <a href="https://zenodo.org/record/7319696#.Y3NuOXbP2Uk">Here</a></p> <p> </p> <p>E: <a href="mailto:molly.fort.21@ucl.ac.uk">molly.fort.21@ucl.ac.uk</a> </p> <p>Object Paradata; </p> <ul> <li><strong>Postcard – c. Early 1900's </strong></li> <li><strong>Language – Eng. </strong></li> <li><strong>Materials – colour print on card, metallic leafing. </strong></li> <li><strong>Front transcription - </strong></li> <li><strong> ‘Greetings’ </strong></li> <li><strong> ‘May your Birthday bring you Peace & perfect Happiness, Golden hopes & Love of Friends, And every Happiness this world can send.’ </strong></li> <li><strong>Object Dimensions – 138mm X 88mm </strong></li> </ul> <p>The postcard is an item of ephemera donated to the UCLDH Digitisation Suite by Prof Melissa Terras, for teaching and training purposes in 2015.</p> <p>This folder contains:</p> <p>Hyperspectral Image data collected using a <a href="https://www.clydehsi.com/hyperspectral-cameras">ClydeHSI VNIR-HR+ Hyperspectral Imaging System</a>.</p> <p>Images captured : 477x484 pixel, 304 spectral band images, 4*4 pixel binning</p> <ul> <li>*.hdr - Header file read out from the ClydeHSI systems instructions for reading the subsequent .raw spectral database. </li> <li>*.raw - Hyperspectral image data cube information. Combination with hdr file creates a ENVI file format, this can be read into a variety of image analysis software packages. </li> <li>postcardhsi.ini - Metadata collected and read out from ClydeHSI system.</li> <li>Dark/White.corr - Correction files taken from camera for processing and minimalising system noise and illumination variences.</li> <li>*_refl.* - Pre - Corrected hyperspectral image data, using provided ClydeHSI software.</li> <li>Truecolour RGB reference image </li> </ul> <p>Each raw and header file set makes-up a single data set in ENVI file format. </p> <p>ENVI reading support exists in Python, R, Matlab, and other common image analysis packages.</p>
Data from "PathOS - D1.2 Scoping Review of Open Science Impact"
<p>This dataset contains the data from the Scoping Review of Open Science Impact. Included are all record for which we assessed the full-text.</p> <p>The columns are as follows:</p> <ul> <li>id: Internal identifier</li> <li>Several metadata columns from Scopus/Web of Science: authors, year, title, abstract, type, DOI</li> <li>OS type: type of Open Science (e.g., Open Access, Citizen Science)</li> <li>inclusion_status: included, duplicate, out of scope, non-english</li> <li>justification: reasons for decision on inclusion_status</li> <li>Several columns with data extracted by the authors: Study details and design, Types of data sources, Study aims, Relevance to which aspect of impact, Key findings, Coverage/Context, Confidence assessment</li> </ul> <p>A complete description of the methods and detailed instructions for coders for extracting data from reports is contained in section 2 of the deliverable report which is available at <a href="https://doi.org/10.5281/zenodo.7883699">https://doi.org/10.5281/zenodo.7883699</a>.</p>
SCShores: time-series of shorelines from Spanish Sandy beaches from citizen-science monitoring program.
<p>This repository contains 5 years of sandy beaches shorelines deriverd from a citizen-science monitoring program in the Spanish coast. The methodology and the dataset are described in:</p> <p><em><strong>González-Villanueva, R., Soriano-González, J., Alejo, I., Criado-Sudau, F., Plomaritis, T., Fernàndez-Mora, À., Benavente, J., Del Río, L., Nombela, M. Á., and Sánchez-García, E.: SCShores: a comprehensive shoreline dataset of Spanish sandy beaches from a citizen-science monitoring programme, Earth System Science Data. V. 15, 4613-4629 , <a href="https://essd.copernicus.org/articles/15/4613/2023/essd-15-4613-2023.html">https://doi.org/10.5194/essd-15-4613-2023</a>, 2023. </strong></em></p> <p>The shoreline dataset is provided in 1 GEOJSON file: SCShores.geojson. This dataset covers five<strong> </strong>sandy beaches located on the Atlantic and Mediterranean coasts of Spain where CoastSnap stations were available, and it includes a total of 1721 shorelines. The coordinate system for the geospatial layer is WGS84.</p> <ul> <li><strong><em>SCShores.geojson</em></strong>: this layer contains the sandy shorelines . Each feature in this layer is a multipoint with the following attributtes: <ul> <li><strong>site</strong>: CoastSnap station name id, e.g. agrelo, samarador, cadiz, ….</li> <li><strong>date</strong>: date and time of the shoreline, yyyyy-mm-dd hh:mm:ss</li> <li><strong>timezone</strong>: Coordinated Universal Time, UTC</li> <li><strong>timestampQuality</strong>: quality flag indicating the confidence in the date-time indicated by the image provider, e.g. 1, 2</li> <li><strong>imageSource</strong>: source of the original image from which the shoreline has been derived, e.g. Instagram, Twitter, Facebook, Email, CoastSnapApp</li> <li><strong>elevation_m:</strong> same as Z coordinate, defined by the observed tide and the tidal offset, in meters, Tide+tide offset</li> <li><strong>verticalDatum</strong>: mean sea level in Alicante, which is considered the zero topographic reference in the Spanish territory, NMMA</li> <li><strong>geometry</strong>: type of geometry used in the file, MultiPoint</li> <li><strong>coordinates</strong>: Geographic WGS84 coordinates for each point in the geometry, longitude, latitude, Z</li> </ul> </li> </ul> <p> </p>
Notably Inaccessible – Data Driven Understanding of Data Science Notebook (In)Accessibility
<p><strong>Overview</strong></p> <p>This dataset artifact contains the intermediate datasets from pipeline executions necessary to reproduce the results of the paper.<br> We share this artifact in hopes of providing a starting point for other researchers to extend the analysis on notebooks, discover more about their accessibility, and offer solutions to make data science more accessible. The scripts needed to generate these datasets and analyse them are shared in the <a href="https://github.com/make4all/notebooka11y">GitHub repository</a> for this work.</p> <blockquote> <p><strong>The dataset contains large files of approximately 60 GB so please exercise caution when extracting the data from compressed files.</strong></p> </blockquote> <blockquote> <p><br> <strong>The dataset contains files which could take a significant amount of run time of the scripts to generate/reproduce.</strong></p> </blockquote> <p><strong>Dataset Contents</strong></p> <p>We briefly summarize the included files in our dataset. Please refer to the <a href="https://github.com/make4all/notebooka11y/blob/main/pipeline/README.md">documentation</a> for specific information about the structure of the data in these files, the scripts to generate them, and runtimes for various parts of our data processing pipeline.</p> <ol> <li><code>epoch_9_loss_0.04706_testAcc_0.96867_X_resnext101_docSeg.pth</code>: We share this model file, originally provided by <a href="https://github.com/jobinkv/DocFigure">Jobin <em>et al.</em></a>, to enable the classification of figures found in our dataset. Please place this into the `model/` <a href="https://github.com/make4all/notebooka11y/tree/main/model">directory</a>.</li> <li><code>model-results.csv</code>: This file contains results from the classification performed on the figures found in the notebooks in our dataset. <blockquote> <p>Performing this classification may take upto a day.</p> </blockquote> </li> <li> <p>a11y-scan-dataset.zip: This archive contains two files and results in datasets of approximately 60GB when extracted. Please ensure that you have sufficient disk space to uncompress this zip archive. The archive contains:</p> <ul> <li> <p><code>a11y/a11y-detailed-result.csv</code>: This dataset contains the accessibility scan results from the scans run on the 100k notebooks across themes.</p> <blockquote><strong>The detailed result file can be really large (> 60 GB) and can be time-consuming to construct.</strong></blockquote> </li> <li> <p><code>a11y/a11y-aggregate-scan.csv</code>: This file is an aggregate of the detailed result that contains the number of each type of error found in each notebook.</p> <blockquote><strong>This file is also shared outside the compressed directory.</strong></blockquote> </li> </ul> </li> <li> <p><code>errors-different-counts-a11y-analyze-errors-summary.csv</code>: This file contains the counts of errors that occur in notebooks across different themes.</p> </li> <li> <p><code>nb_processed_cell_html.csv</code>: This file contains metadata corresponding to each cell extracted from the html exports of our notebooks.</p> </li> <li> <p><code>nb_first_interactive_cell.csv</code>: This file contains the necessary metadata to compute the first interactive element, as defined in our paper, in each notebook.</p> </li> <li> <p><code>nb_processed.csv</code>: This file contains the necessary data after processing the notebooks extracting the number of images, imports, languages, and cell level information.</p> </li> <li> <p><code>processed_function_calls.csv</code>: This file contains the information about the notebooks, the various imports and function calls used within the notebooks.</p> </li> </ol>
Data from “A Mixed Method Approach to Understanding the Public Health Impact of a School-Based Citizen Science Program to Reduce Arsenic in Private Well Water”
Objectives We have approached the problem of low well water testing rates in Maine and New Hampshire communities by developing the All About Arsenic (AAA) project, which engages secondary school teachers and students as citizen scientists in collecting well water samples for analysis of arsenic and other toxic metals and supports their outreach efforts to their communities. Methods We assessed this project’s public health impact by analyzing student data relative to existing well water quality datasets in both states. In addition, we surveyed private well owners who contributed well water samples to the project to determine the actions taken to mitigate arsenic in well water. Data The data presented here are used in the analyses performed for the publication: "A Mixed Method Approach to Understanding the Public Health Impact of a School-Based Citizen Science Program to Reduce Arsenic in Private Well Water.” Additional data may be available at: The Anecdata Project Page: https://anecdata.org/projects/view/299 The project website: https://www.allaboutarsenic.org/
The recovery of plant community composition following passive restoration across spatial scales, Cedar Creek Ecosystem Science Reserve, 1983-2016
1. Human impacts have led to dramatic biodiversity change which can be highly scale-dependent across space and time. A primary means to manage these changes is via passive (here, the removal of disturbance) or active (management interventions) ecological restoration. The recovery of biodiversity, following the removal of disturbance is often incomplete relative to some kind of reference target. The magnitude of recovery of ecological systems following disturbance depend on the landscape matrix, as well as the temporal and spatial scales at which biodiversity is measured. 2. We measured the recovery of biodiversity and species composition over 27 years in 17 temperate grasslands abandoned after agriculture at different points in time, collectively forming a chronosequence since abandonment from one to eighty years. We compare these abandoned sites with known agricultural land-use histories to never-disturbed sites as relative benchmarks. We specifically measured aspects of diversity at the local plot-scale (α-scale, 0.5m2) and site-scale (γ-scale, 10m2), as well as the within-site heterogeneity (β-diversity) and among-site variation in species composition (turnover and nestedness). 3. At our α-scale, sites recovering after agricultural abandonment only had 70% of the plant species richness (and ~30% of the evenness), compared to never-ploughed sites. Within-site β-diversity recovered following agricultural abandonment to around 90% after 80 years. This effect, however, was not enough to lead to recovery at our γ-scale. Richness in recovering sites was ~65% of that in remnant never-ploughed sites. The presence of species characteristic of the never disturbed sites increased in the recovering sites through time. Forb and legume cover declines in years since abandonment, relative to graminoid cover across sites. 4. Synthesis. We found that, during the 80 years after agricultural abandonment, old-fields did not recover to the level of biodiversity in remnant never-plough
Temperature logger deployment methods and irradiance-biased temperature data, King Abdullah University of Science and Technology, Red Sea, 2023.
Solar irradiance can offset the temperature recorded by underwater sensing instruments (aka "loggers"). We collected temperature and PAR (photosynthetic active radiation) data during two short-term in situ deployments on a shallow fringing reef adjacent to the King Abdullah University of Science and Technology (KAUST) in the Red Sea. The first deployment quantified the measurement bias due to solar heating over five days in February 2023 while the second compared the effect of different shading methods on logger performance over 24 hours in June 2023. We also recorded temperature in a controlled calibration bath in the lab with ten of the most widely used loggers to further assess their accuracy, response time, and intra-logger variation. Finally, to understand current practices of measuring temperature on coral reefs, we summarized logger deployment method details from a literature review of coral reef studies published from 2013 to 2022. Such details included how often loggers recorded the temperature, the depth where loggers were deployed, and whether the authors reported shading or protecting their loggers. This data package is complete and part of a larger project that aims to develop an instrument deployment framework for restoration-based reef monitoring, which includes instrument recommendations and deployment guidelines.
Course Materials for Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730)
In today's world, understanding environmental data and making informed decisions based on it is crucial for addressing complex environmental challenges. Yale School of the Environment's Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730) course serves as an introduction to the integration of environmental data using R programming language, coupled with machine learning techniques. This dataset contains a zip file with all the data files used in this course, along with a README that has the metadata for those files.
MCR LTER: Coral Reef: Biodiversity has a positive but saturating effect on imperiled coral reefs; data for Clements and Hay 2021, Science Advances
Species loss threatens ecosystems worldwide, but the ecological processes and thresholds that underpin positive biodiversity effects among critically important foundation species, such as corals on tropical reefs, remain inadequately understood. In field experiments, we manipulated coral species richness and intraspecific density to test whether, and how, biodiversity affects coral productivity and survival. Corals performed better in mixed species assemblages. Improved performance was unexplained by competition theory alone, suggesting that positive effects exceeded agonistic interactions during our experiments. Peak coral performance occurred at intermediate species richness and declined thereafter. Positive effects of coral diversity suggest that species’ losses on degraded reefs make recovery more difficult and further decline more likely. Harnessing these positive interactions may improve ecosystem conservation and restoration in a changing ocean. This material is based upon work supported by the U.S. National Science Foundation under Grant No. OCE 16-37396 (and earlier awards) as well as a generous gift from the Gordon and Betty Moore Foundation. Research was completed under permits issued by the French Polynesian Government (Délégation à la Recherche) and the Haut-commissariat de la République en Polynésie Francaise (DTRT) (Protocole d'Accueil 2005-2022). This work represents a contribution of the Moorea Coral Reef (MCR) LTER Site. Datasets used in this study are available online from the BCO-DMO data system. Data for this paper can be found at (https://www.bco-dmo.org/project/837802).
Open Science: New Challenges and Opportunities for the PV sector
<p>Presentation given at the European PV Solar Energy Conference, Marseille, 2019 about the development of Open Science in the context of photovoltaics</p>
A Dataset of Science-Policy Collaborations from the International Geneva Ecosystem
<p><strong>This dataset contains data on the characteristics of science-policy collaborations (SPCs) as well as on the perceived value of the Geneva Science-Policy Interface as an intermediary actor facilitating collaborative processes. </strong></p> <p>----</p> <p>In 2020, the Geneva Science-Policy Interface deployed a call for projects explicitly targeting SPCs. The purpose of the call for projects is to motivate policy actors and scientists to join their efforts and propose collaboration projects, select the most promising projects, and provide them with financial and strategic support. Projects are eligible only if both academia and policy actors are represented, and if policy actors are part of the International Geneva ecosystem.</p> <p>The call for projects is a structured and standardised process that allows for the collection of identical data across a sample of SPCs, thus enabling comparative analyses. Data is collected through crowdsourcing, coupling challenge-based data collection and survey data. The Geneva Science-Policy Interface uses the call for projects to collect data on characteristics of SPCs as well as on the perceived value of the GSPI as an intermediary actor facilitating collaborative processes. </p>
Practices and policies of preprint platforms for life and biomedical sciences
<p>Given the increase in the use and profile of preprint servers – and alternative publishing hybrid platforms such as F1000 Research – in the life sciences, it is increasingly important to identify how many such servers and hybrids exist, to describe their scope in terms of the scientific disciplines they cover, and to compare and contrast their characteristics and policies.</p> <p>We surveyed forty-four (44) platforms that host preprints relevant to life and biomedical sciences and that were active online and accepting submissions on 25 June 2019. Information on preprint platform policies, features and practices was collected through online research by the authors and by surveying preprint platform representatives directly. </p> <p>Full data sheets include an additional 5 platforms hosted on OSF Preprints (rows 49-53) to fulfil the wider scope for the ASAPbio project, not in disciplinary scope (biology and medical sciences) for the manuscript with Jamie Kirkham.</p> <p><strong>Tables 1-5: </strong>Data (44 platforms, manuscript) are separated into five main tables of information and a list of preprint platform websites for reference.</p> <p>Table 1: Scope and ownership of each server<br> Table 2: Content-specific characteristics and information relating to submission, journal transfer options, and external discoverability<br> Table 3: Screening, moderation, and permanence of content<br> Table 4: Usage metrics and other features<br> Table 5: Metadata<br> Preprint platform websites</p> <p>Data for each platform are listed as ‘Verified’ in the tables if these tables (V1.0 or V2.0) were seen and approved by a platform representative between January 13 and January 27, 2020.</p> <p><strong>Original online survey:</strong> a blank copy of the original survey form used by online researchers (the authors) and supplied pre-filled (or empty, in some cases) to preprint platform representatives for verification (or completion, in some cases). </p> <p><strong>Final data:</strong> survey data is presented in .txt and .xlsx, as follows:</p> <ul> <li>Row 1: Heading (where field is included in manuscript tables, the heading presented here replaces any heading used in original survey. All columns are presented in the order the information was requested on the original survey form, with some supplementary columns added and columns removed (detailed below).</li> <li>Row 2: Schema or description of field</li> <li>Row 3: Whether and where included in manuscript tables. For supporting information for table data (e.g. source information, URLs), the table location for supported data is indicated in brackets, e.g. (Table 2) and supporting information is not included in tables. Data included in manuscript tables is presented in its final form, which in some cases is simplified from the original survey data. This simplified version of the data was presented to platform representatives for additional verification (v1.0/v2.0 verification). Data not included in manuscript tables is presented here as verified by platform representatives and/or found online. Some columns from the original survey have been removed due to the information not being informative or useful: specifically, Print ISSN (not reported for any platform); End date (no platforms have an end date; although two platforms stopped accepting submissions after survey completed; Personal contact information for platform representative(s) has been removed).</li> <li>Rows 4 onwards: data for each preprint platform (44 included in manuscript (rows 4-47), plus 5 additional OSF platforms (rows 48-52)</li> <li>Columns 3-6 (D-G) report online research and verification information and Column 13 (M) reports an additional data field (number of articles) – these are supplementary to the original survey columns</li> <li>Verification status: Released V1/V2 data applies to data included in manuscript tables (as indicated in row 3); Online survey data applies to data used for manuscript tables and also to original survey data included here but not included in manuscript tables (‘Not included’ in row 3)</li> <li>Note that data fields are presented as individual columns in these sheets, while some entries in Tables 1-5 combine several data fields.</li> </ul> <p>These data were collected in collaboration and as part of:<br> i. An ASAPbio project, led by Dr Naomi Penfold, to develop an online directory of preprint platforms<br> ii. A research project led by Prof Jamie Kirkham<br> These data are supplementary outputs for both projects.</p> <p>Data v1.0 were presented during the ASAPbio January 2020 workshop – see Penfold, Naomi C, & Polka, Jessica. (2020, January). ASAPbio Preprint Platform Directory: 2019 data (presentation) (Version 1.0). Zenodo. http://doi.org/10.5281/zenodo.3626770.<br> <br> <strong>Version 3.0 updates (December 14, 2020): added new files with updated information about servers from the ASAPbio preprint directory (https://asapbio.org/preprint-servers), provided by Jessica Polka (now included as author).</strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.