Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
773
datasets available to search
ShareScore release 0.7.1
Dataset results
773 results for “data science”
Notably Inaccessible – Data Driven Understanding of Data Science Notebook (In)Accessibility
<p><strong>Overview</strong></p> <p>This dataset artifact contains the intermediate datasets from pipeline executions necessary to reproduce the results of the paper.<br> We share this artifact in hopes of providing a starting point for other researchers to extend the analysis on notebooks, discover more about their accessibility, and offer solutions to make data science more accessible. The scripts needed to generate these datasets and analyse them are shared in the <a href="https://github.com/make4all/notebooka11y">GitHub repository</a> for this work.</p> <blockquote> <p><strong>The dataset contains large files of approximately 60 GB so please exercise caution when extracting the data from compressed files.</strong></p> </blockquote> <blockquote> <p><br> <strong>The dataset contains files which could take a significant amount of run time of the scripts to generate/reproduce.</strong></p> </blockquote> <p><strong>Dataset Contents</strong></p> <p>We briefly summarize the included files in our dataset. Please refer to the <a href="https://github.com/make4all/notebooka11y/blob/main/pipeline/README.md">documentation</a> for specific information about the structure of the data in these files, the scripts to generate them, and runtimes for various parts of our data processing pipeline.</p> <ol> <li><code>epoch_9_loss_0.04706_testAcc_0.96867_X_resnext101_docSeg.pth</code>: We share this model file, originally provided by <a href="https://github.com/jobinkv/DocFigure">Jobin <em>et al.</em></a>, to enable the classification of figures found in our dataset. Please place this into the `model/` <a href="https://github.com/make4all/notebooka11y/tree/main/model">directory</a>.</li> <li><code>model-results.csv</code>: This file contains results from the classification performed on the figures found in the notebooks in our dataset. <blockquote> <p>Performing this classification may take upto a day.</p> </blockquote> </li> <li> <p>a11y-scan-dataset.zip: This archive contains two files and results in datasets of approximately 60GB when extracted. Please ensure that you have sufficient disk space to uncompress this zip archive. The archive contains:</p> <ul> <li> <p><code>a11y/a11y-detailed-result.csv</code>: This dataset contains the accessibility scan results from the scans run on the 100k notebooks across themes.</p> <blockquote><strong>The detailed result file can be really large (> 60 GB) and can be time-consuming to construct.</strong></blockquote> </li> <li> <p><code>a11y/a11y-aggregate-scan.csv</code>: This file is an aggregate of the detailed result that contains the number of each type of error found in each notebook.</p> <blockquote><strong>This file is also shared outside the compressed directory.</strong></blockquote> </li> </ul> </li> <li> <p><code>errors-different-counts-a11y-analyze-errors-summary.csv</code>: This file contains the counts of errors that occur in notebooks across different themes.</p> </li> <li> <p><code>nb_processed_cell_html.csv</code>: This file contains metadata corresponding to each cell extracted from the html exports of our notebooks.</p> </li> <li> <p><code>nb_first_interactive_cell.csv</code>: This file contains the necessary metadata to compute the first interactive element, as defined in our paper, in each notebook.</p> </li> <li> <p><code>nb_processed.csv</code>: This file contains the necessary data after processing the notebooks extracting the number of images, imports, languages, and cell level information.</p> </li> <li> <p><code>processed_function_calls.csv</code>: This file contains the information about the notebooks, the various imports and function calls used within the notebooks.</p> </li> </ol>
Data from “A Mixed Method Approach to Understanding the Public Health Impact of a School-Based Citizen Science Program to Reduce Arsenic in Private Well Water”
Objectives We have approached the problem of low well water testing rates in Maine and New Hampshire communities by developing the All About Arsenic (AAA) project, which engages secondary school teachers and students as citizen scientists in collecting well water samples for analysis of arsenic and other toxic metals and supports their outreach efforts to their communities. Methods We assessed this project’s public health impact by analyzing student data relative to existing well water quality datasets in both states. In addition, we surveyed private well owners who contributed well water samples to the project to determine the actions taken to mitigate arsenic in well water. Data The data presented here are used in the analyses performed for the publication: "A Mixed Method Approach to Understanding the Public Health Impact of a School-Based Citizen Science Program to Reduce Arsenic in Private Well Water.” Additional data may be available at: The Anecdata Project Page: https://anecdata.org/projects/view/299 The project website: https://www.allaboutarsenic.org/
Temperature logger deployment methods and irradiance-biased temperature data, King Abdullah University of Science and Technology, Red Sea, 2023.
Solar irradiance can offset the temperature recorded by underwater sensing instruments (aka "loggers"). We collected temperature and PAR (photosynthetic active radiation) data during two short-term in situ deployments on a shallow fringing reef adjacent to the King Abdullah University of Science and Technology (KAUST) in the Red Sea. The first deployment quantified the measurement bias due to solar heating over five days in February 2023 while the second compared the effect of different shading methods on logger performance over 24 hours in June 2023. We also recorded temperature in a controlled calibration bath in the lab with ten of the most widely used loggers to further assess their accuracy, response time, and intra-logger variation. Finally, to understand current practices of measuring temperature on coral reefs, we summarized logger deployment method details from a literature review of coral reef studies published from 2013 to 2022. Such details included how often loggers recorded the temperature, the depth where loggers were deployed, and whether the authors reported shading or protecting their loggers. This data package is complete and part of a larger project that aims to develop an instrument deployment framework for restoration-based reef monitoring, which includes instrument recommendations and deployment guidelines.
Course Materials for Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730)
In today's world, understanding environmental data and making informed decisions based on it is crucial for addressing complex environmental challenges. Yale School of the Environment's Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730) course serves as an introduction to the integration of environmental data using R programming language, coupled with machine learning techniques. This dataset contains a zip file with all the data files used in this course, along with a README that has the metadata for those files.
MCR LTER: Coral Reef: Biodiversity has a positive but saturating effect on imperiled coral reefs; data for Clements and Hay 2021, Science Advances
Species loss threatens ecosystems worldwide, but the ecological processes and thresholds that underpin positive biodiversity effects among critically important foundation species, such as corals on tropical reefs, remain inadequately understood. In field experiments, we manipulated coral species richness and intraspecific density to test whether, and how, biodiversity affects coral productivity and survival. Corals performed better in mixed species assemblages. Improved performance was unexplained by competition theory alone, suggesting that positive effects exceeded agonistic interactions during our experiments. Peak coral performance occurred at intermediate species richness and declined thereafter. Positive effects of coral diversity suggest that species’ losses on degraded reefs make recovery more difficult and further decline more likely. Harnessing these positive interactions may improve ecosystem conservation and restoration in a changing ocean. This material is based upon work supported by the U.S. National Science Foundation under Grant No. OCE 16-37396 (and earlier awards) as well as a generous gift from the Gordon and Betty Moore Foundation. Research was completed under permits issued by the French Polynesian Government (Délégation à la Recherche) and the Haut-commissariat de la République en Polynésie Francaise (DTRT) (Protocole d'Accueil 2005-2022). This work represents a contribution of the Moorea Coral Reef (MCR) LTER Site. Datasets used in this study are available online from the BCO-DMO data system. Data for this paper can be found at (https://www.bco-dmo.org/project/837802).
Identification at local and global scale: a case for using the Compact URI (CURIE) for life science data
<p>Panel A) A Local Resource Identifier (LRI) is not suited to global scale identification because of inevitable collisions: “9606” corresponds to a Pubmed article, a CGNC gene, a PubChem chemical, as well as an NCBI taxon (<em>Homo sapiens</em>), a BOLD taxon (<em>Bombycilla</em> <em>cedrorum</em>), and a GRIN taxon (<em>Catha</em> <em>edulis</em>)</p> <p>Panel B) Prefixing is often used to indicate the source of an LRI, but prefixes themselves are often undocumented and collide.</p> <p>Panel C) Prefixes may exist in alternate forms. When all of the alternates are not known, collapsing equivalent identifiers is tedious and incomplete.</p> <p>Panel D) CURIE syntax addresses these issues by having a prefix whose relationship with a resolving namespace is clearly documented.</p>
Video 5 - Open Science: why do we need data stewards.
<p><span>An interview on the need of data professionals and Open Science skills with York Sure-Vetter, Director of NFDI, Germany and Professor at Karlsruhe Institute of Technology; Jessica Lindvall, Head of Training at SciLifeLab Training Hub; Anne Sophie Fink, Head of Data Management at DeiC (Denmark); and Sally Chambers, Director at DARIAH-EU.</span></p> <p><span>Modern research and technology can require not only large amounts of data, but also good data quality. Ensuring good data quality requires specialized expertise. Data stewards and other data management professionals support researchers with providing and working with good quality data which are also FAIR. Open software, open infrastructures are also important bits in the Open Science puzzle, all of which require funding. The uptake of Open Science depends on widespread Open Science awareness and skills. These require outreach, training and formal education.</span></p>
Data and code associated with "The Observed Availability of Data and Code in Earth Science and Artificial Intelligence"
<p>Data and code associated with "The Observed Availability of Data and Code in Earth Science <br>and Artificial Intelligence" by Erin A. Jones, Brandon McClung, Hadi Fawad, and Amy McGovern.</p> <p>Instructions: To reproduce figures, download all associated Python and CSV files and place<br> in a single directory.<br> Run BAMS_plot.py as you would run Python code on your system.</p> <p>Code:<br>BAMS_plot.py: Python code for categorizing data availability statements based on given data<br> documented below and creating figures 1-3. </p> <p> Code was originally developed for Python 3.11.7 and run in the Spyder <br> (version 5.4.3) IDE.<br> <br> Libraries utilized:<br> numpy (version 1.26.4) <br> pandas (version 2.1.4)<br> matplotlib (version 3.8.0)<br> <br> For additional documentation, please see code file.</p> <p>Data:<br>ASDC_AIES.csv: CSV file containing relevant availability statement data for Artificial <br> Intelligence for the Earth Systems (AIES)<br>ASDC_AI_in_Geo.csv: CSV file containing relevant availability statement data for Artificial <br> Intelligence in Geosciences (AI in Geo.)<br>ASDC_AIJ.csv: CSV file containing relevant availability statement data for Artificial <br> Intelligence (AIJ)<br>ASDC_MWR.csv: CSV file containing relevant availability statement data for Monthly <br> Weather Review (MWR)<br><br></p> <p><br>Data documentation:<br>All CSV files contain the same format of information for each journal. The CSV files above are <br>needed for the BAMS_plot.py code attached.</p> <p>Records were analyzed based on the criteria below.</p> <p> Records:<br> 1) Title of paper<br> The title of the examined journal article.<br> 2) Article DOI (or URL)<br> A link to the examined journal article. For AIES, AI in Geo., MWR, the DOI is <br> generally given. For AIJ, the URL is given.<br> 3) Journal name<br> The name of the journal where the examined article is published. Either a full<br> journal name (e.g., Monthly Weather Review), or the acronym used in the <br> associated paper (e.g., AIES) is used.<br> 4) Year of publication<br> The year the article was posted online/in print.<br> 5) Is there an ASDC?<br> If the article contains an availability statement in any form, "yes" is <br> recorded. Otherwise, "no" is recorded.<br> 6) Justification for non-open data?<br> If an availability statement contains some justification for why data is not <br> openly available, the justification is summarized and recorded as one of the <br> following options: 1) Dataset too large, 2) Licensing/Proprietary, 3) Can be <br> obtained from other entities, 4) Sensitive information, 5) Available at later <br> date. If the statement indicates any data is not openly available and no <br> justification is provided, or if no statement is provided is provided "None" <br> is recorded. If the statement indicates openly available data or no data <br> produced, "N/A" is recorded.<br> 7) All data available<br> If there is an availability statement and data is produced, "y" is recorded <br> if means to access data associated with the article are given and there is no <br> indication that any data is not openly available; "n" is recorded if no means <br> to access data are given or there is some indication that some or all data is <br> not openly available. If there is no availability statement or no data is <br> produced, the record is left blank.<br> 8) At least some data available<br> If there is an availability statement and data is produced, "y" is recorded <br> if any means to access data associated with the article are given; "n" is <br> recorded if no means to access data are given. If there is no availability <br> statement or no data is produced, the record is left blank.<br> 9) All code available<br> If there is an availability statement and data is produced, "y" is recorded <br> if means to access code associated with the article are given and there is no <br> indication that any code is not openly available; "n" is recorded if no means <br> to access code are given or there is some indication that some or all code is <br> not openly available. If there is no availability statement or no data is <br> produced, the record is left blank.<br> 10) At least some code available<br> If there is an availability statement and data is produced, "y" is recorded <br> if any means to access code associated with the article are given; "n" is <br> recorded if no means to access code are given. If there is no availability <br> statement or no data is produced, the record is left blank.<br> 11) All data available upon request<br> If there is an availability statement indicating data is produced and no data <br> is openly available, "y" is recorded if any data is available upon request to <br> the authors of the examined journal article (not a request to any other <br> entity); "n" is recorded if no data is available upon request to the authors <br> of the examined journal article. If there is no availability statement, any <br> data is openly available, or no data is produced, the record is left blank.<br> 12) At least some data available upon request<br> If there is an availability statement indicating data is produced and not all <br> data is openly available, "y" is recorded if all data is available upon <br> request to the authors of the examined journal article (not a request to any <br> other entity); "n" is recorded if not all data is available upon request to <br> the authors of the examined journal article. If there is no availability <br> statement, all data is openly available, or no data is produced, the record<br> is left blank.<br> 13) no data produced<br> If there is an availability statement that indicates that no data was<br> produced for the examined journal article, "y" is recorded. Otherwise, the<br> record is left blank.<br> 14) links work<br> If the availability statement contains one or more links to a data or code <br> repository, "y" is recorded if all links work; "n" is recorded if one or more <br> links do not work. If there is no availability statement or the statement <br> does not contain any links to a data or code repository, the record is left <br> blank. </p>
Data Science job offers in Euraxess.
<p>It is not always easy to find job opportunities if you are interested in beginning to do research in a certain field. In this sense, having an up-to-date dataset with job offers in your field of interest would simplify this search. This dataset could be generated using web scraping methods.</p> <p>Although the web scraper we built could be applied to every field, in this project we focused in opportunities related with data science (i.e. data scientist, data analyst, data engineer...) published on <a href="https://euraxess.ec.europa.eu/">EURAXESS</a>.</p> <p>The dataset generated with this package contains job offers obtained from EURAXESS. Each row of the dataset contains different job offers and its attributes. In the example table showed below, the dataset was obtained using "Data Scientist" as keyword, but another keywords would result in different datasets. The columns describing the dataset are:</p> <ul> <li>Job Offer Title: Title of the job offer.</li> <li>Researcher Profile: Expected applicant profile/s.</li> <li>Company: Company offering the job.</li> <li>Hours/Week: Weekly working hours.</li> <li>Country: Country where the job is offered.</li> <li>City: City where the job is offered.</li> <li>Where to Apply: Url or email where to apply to the offer.</li> <li>More info: URL where the offer can be located.</li> </ul> <p>Dataset generated by web scraping methods: https://github.com/avicenteg/euraxess_scraping</p>
Dataset for Paper "Towards Increased Diversity in STEM Education: Five archetypes Derived through a Data-Driven Approach Examining a Computer Science Student Cohort
<p># Dataset for Paper "Towards Increased Diversity in STEM Education: Five archetypes Derived through a Data-Driven Approach Examining a Computer Science Student Cohort" - Rev #1</p> <p>This is the dataset for the paper titled "Towards Increased Diversity in STEM Education: Five archetypes Derived through a Data-Driven Approach Examining a Computer Science Student Cohort".</p> <p>In case of questions, feel free to contact the authors, *anonymised*, ORCID: https://orcid.org/*anonymised*, current affiliation and email: *anonymised*</p> <p>## Survey 2019 ##<br> The raw survey data for the initial 2019 survey is available in the file *survey2019_anon.csv*. Note that the data is anonymised as free-text comments have been removed. Explanations on the variables and their levels are given in the files *variables_survey2019.csv* and *values_survey2019.csv*.<br> The questionnaire for the 2019 survey is contained in *survey2019_instrument.pdf*.</p> <p>## Survey 2020 ##<br> The raw survey data for the 2020 survey is available in the file *rdata_anon_survey2020.csv*. Additional scripts are supplied to reproduce the exploratory factor analysis. The main entry is the file *EFA.R*, which imports the data. The file contains some comments on the process.<br> The questionnaire for the 2020 survey is contained in *survey2020_instrument.pdf*.</p> <p>## Interviews ##<br> The interview guide used for the five interviews is available in the file *interview_instrument.pdf*.</p>
Data/ codes used in the the Natural Hazards and Earth System Sciences (NHESS) publication titled "Wind-Wave Characteristics and extremes along the Emilia-Romagna coast" by Pranavam Ayyappan Pillai et al. (2022)
<p>The archive contains datasets and codes used in the manuscript titled "Wind-Wave Characteristics and extremes along the Emilia-Romagna coast", and published in the journal <em>Natural Hazards and Earth System Sciences</em> (<em>NHESS</em>) by Pranavam Ayyappan Pillai et al., 2022.</p> <p>Pranavam Ayyappan Pillai, U., Pinardi, N., Federico, I., Causio, S., Trotta, F., Unguendoli, S., and Valentini, A.: Wind-Wave Characteristics and extremes along the Emilia-Romagna coast, Nat. Hazards Earth Syst. Sci. Discuss. https://doi.org/10.5194/nhess-2022-103, 2022.</p>
Biological data science courses at UMONS, Belgium: student's activity for 2019-2020
<p>Progression of the students in the different exercises of the biological data science courses at the University of Mons, Belgium for the academic year 2019-2020.</p> <p>Activity of the students was recorded to monitor their individual progression in asynchronous exercises. The courses were taught in flipped classroom by Philippe Grosjean (<a href="mailto:philippe.grosjean@umons.ac.be">philippe.grosjean@umons.ac.be</a>) and Guyliann Engels (<a href="mailto:guyliann.engels@umons.ac.be">guyliann.engels@umons.ac.be</a>) the University of Mons. These authors designed almost all the teaching material, the exercises, and the related software. The courses were also taught at the Campus Charleroi by Raphaël Conotte (<a href="mailto:raphael.conotte@umons.ac.be">raphael.conotte@umons.ac.be</a>) that also contributed to a part of the learnr exercises and of the inline course.</p> <p><strong>How to use these data?</strong></p> <p>The README file provides detailed information on the purpose, collection and management of the data. The data are presented in tabular format in CSV files. Metadata in the `datapackage.json` document the different tables and their fields. It is in the Frictionless data format (<a href="https://frictionlessdata.io/">https://frictionlessdata.io</a>). You can get a view of a part of these metadata by uploading the file `datapackage.json` into the inline data package creator at <a href="https://create.frictionlessdata.io/">https://create.frictionlessdata.io</a>. There is a large set of libraries and tools for different programming languages available at <a href="https://frictionlessdata.io/tooling/libraries/">https://frictionlessdata.io/tooling/libraries/</a>. Otherwise, any CSV library should import the data in your favourite software. Please, note that encoding is UTF8. For R, the {learnitdown} package provides specific functions to import these data and/or convert them in a SQLite database (<a href="https://www.sciviews.org/learnitdown/">https://www.sciviews.org/learnitdown/</a>).</p> <p>For any question, send an email at <a href="mailto:sdd@sciviews.org">sdd@sciviews.org</a>.</p>
Biological data science courses at UMONS, Belgium: student's activity for 2020-2021
<p>Progression of the students in the different exercises of the biological data science courses at the University of Mons, Belgium for the academic year 2020-2021.</p> <p>Activity of the students was recorded to monitor their individual progression in asynchronous exercises. The courses were taught in flipped classroom by Philippe Grosjean (<a href="mailto:philippe.grosjean@umons.ac.be">philippe.grosjean@umons.ac.be</a>) and Guyliann Engels (<a href="mailto:guyliann.engels@umons.ac.be">guyliann.engels@umons.ac.be</a>) the University of Mons. These authors designed almost all the teaching material, the exercises, and the related software. The courses were also taught at the Campus Charleroi by Raphaël Conotte (<a href="mailto:raphael.conotte@umons.ac.be">raphael.conotte@umons.ac.be</a>) that also contributed to a part of the learnr exercises and of the inline course.</p> <p><strong>How to use these data?</strong></p> <p>The README file provides detailed information on the purpose, collection and management of the data. The data are presented in tabular format in CSV files. Metadata in the `datapackage.json` document the different tables and their fields. It is in the Frictionless data format (<a href="https://frictionlessdata.io">https://frictionlessdata.io</a>). You can get a view of a part of these metadata by uploading the file `datapackage.json` into the inline data package creator at <a href="https://create.frictionlessdata.io">https://create.frictionlessdata.io</a>. There is a large set of libraries and tools for different programming languages available at <a href="https://frictionlessdata.io/tooling/libraries/">https://frictionlessdata.io/tooling/libraries/</a>. Otherwise, any CSV library should import the data in your favourite software. Please, note that encoding is UTF8. For R, the {learnitdown} package provides specific functions to import these data and/or convert them in a SQLite database (<a href="https://www.sciviews.org/learnitdown/">https://www.sciviews.org/learnitdown/</a>).</p> <p>For any question, send an email at <a href="mailto:sdd@sciviews.org">sdd@sciviews.org</a>.</p>
Biological data science courses at UMONS, Belgium: student's activity for 2018-2019
<p>Progression of the students in the different exercises of the biological data science courses at the University of Mons, Belgium for the academic year 2018-2019.</p> <p>Activity of the students was recorded to monitor their individual progression in asynchronous exercises. The courses were taught in flipped classroom by Philippe Grosjean (<a href="mailto:philippe.grosjean@umons.ac.be">philippe.grosjean@umons.ac.be</a>) and Guyliann Engels (<a href="mailto:guyliann.engels@umons.ac.be">guyliann.engels@umons.ac.be</a>) the University of Mons. These authors designed almost all the teaching material, the exercises, and the related software.</p> <p><strong>How to use these data?</strong></p> <p>The README file provides detailed information on the purpose, collection and management of the data. The data are presented in tabular format in CSV files. Metadata in the `datapackage.json` document the different tables and their fields. It is in the Frictionless data format (<a href="https://frictionlessdata.io/">https://frictionlessdata.io</a>). You can get a view of a part of these metadata by uploading the file `datapackage.json` into the inline data package creator at <a href="https://create.frictionlessdata.io/">https://create.frictionlessdata.io</a>. There is a large set of libraries and tools for different programming languages available at <a href="https://frictionlessdata.io/tooling/libraries/">https://frictionlessdata.io/tooling/libraries/</a>. Otherwise, any CSV library should import the data in your favourite software. Please, note that encoding is UTF8. For R, the {learnitdown} package provides specific functions to import these data and/or convert them in a SQLite database (<a href="https://www.sciviews.org/learnitdown/">https://www.sciviews.org/learnitdown/</a>).</p> <p>For any question, send an email at <a href="mailto:sdd@sciviews.org">sdd@sciviews.org</a>.</p>
Code and data accompanying Palmeirim et al. (2022) Emergent properties of species-habitat networks in an insular forest landscape. Science Advances
<p>Dataset containing species distribution in insular forest fragments at Balbina and full R code for analyses and figures.</p> <p>For deatails, please see the original publication: "Emergent properties of species-habitat networks in an insular forest landscape". Ana Filipa Palmeirim, Carine Emer, Maíra Benchimol, Danielle Storck-Tonon, Anderson S. Bueno, Carlos A. Peres. Science Advances (2022). 10.1126/sciadv.abm0397.</p> <p> </p>
Unprocessed data from the Jungle Weather Zooniverse citizen science project
<p>The <a href="https://www.zooniverse.org/projects/khufkens/jungle-weather">Jungle Weather project</a> aimed to transcribe weather observations recorded between 1949 and 1958 in the tropical rainforest of the Democratic Republic of the Congo. Long-term observations of tropical weather are rare. The Jungle Weather, as part of the COBECORE project, contains observations of three decades of data of weather in the central African tropical forest, and are therefore an extraordinary source of information to support our understanding of for example drought resilience of trees species.</p> <p><strong>Summary</strong></p> <p>Both input and output of the citizen science transcriptions are provided in this data set. This includes the original cut-outs as used in the Zoonivese project, and the output as generated by the Zooniverse data export routines. The data export routines provided CSV output with JSON subfields on the content of each classification made. In addition, we provided the exported subject list and the details of each workflow.</p> <p>In total the project output constitutes of four files:</p> <ul> <li>transcribe-climate-data-classifications.csv (annotations of the table cells)</li> <li>transcribe-meta-data-classifications.csv (annotations of table headers)</li> <li>jungle-weather-workflows.csv (description of the citsci workflow)</li> <li>jungle-weather-subjects.csv (list of all images transcribed, and their online location for validation / referencing)</li> </ul> <p>and roughly ~3GB in data volume.</p> <p><strong>Context</strong></p> <p>Our understanding of forest ecosystem responses to climate change relies on consistent long-term observations to provide baseline measurements. In the central Congo Basin established long-term observation programs are rare. In terms of meteorological observations, the central Congo Basin is currently represented by only a few rain gauges, limiting climate forecasts across the Congo Basin and the central African continent. This lack of long-term (historical) climatological data leaves the central Congo Basin spatially and temporally under-represented. However, old climate records could provide valuable information about previous growing conditions of the forest.</p> <p>Large amounts of ecological and climatological data, approximately five decades (~1910 – 1960), exists as unexplored heritage, stored in various Belgian federal archives and collections. As part of a larger project called Congo Basin eco-climatological data recovery and valorization (COBECORE, see) the Jungle Weather project will help transcribe historical climatological data as measured throughout the Congo Basin. These data will in part complement the completed <a href="https://www.zooniverse.org/projects/khufkens/jungle-rhythms">Jungle Rhythms Zooniverse project</a>, further valorizing these transcribed data.</p> <p><strong>Historical data</strong></p> <p>Within this project we will focus on data records as recorded throughout the tropical part of what is currently the Democratic Republic of the Congo (DRC). The area which we will cover is shown above in the map as an open polygon. The project will not cover the southern province of Katanga (red crosshatches) as this area transitions here from tropical to a humid subtropical climate.</p> <p>The historical data is archived and stored in the Belgian State Archives. The Belgian State Archive harbour almost all data regarding colonial affairs, ranging from communications about trade to the raw data as digitized within the context of the Jungle Weathers project. Row upon row of data is stored in the basement. Below you see a part of the INEAC (Institut National pour l’Etude Agronomique du Congo belge) archive, which holds all climatological records.</p> <p>These climatological records were noted rigorously on carbon copy paper. However, due to the hand written nature of the data (and the volume involved) automated processing is not possible. Although optical character recognition (OCR) works wonderfully on printed data the high variability in characters and the low contrast pencil markings contribute to the failure of current automated approaches. Similar to the <a href="https://www.oldweather.org/">Old Weather project</a> and in spirit of the Jungle Rhythms project, a keen eye is required to decipher the numbers written down on these sheets.</p> <p><strong>Pre-processing / digitization</strong></p> <p>The project provided citizen scientists with digital pictures of the original sheets. Scanning these climate data sheets was a laborious process. In total more than 70 000 records were digitized. Unlike the Old Weather project we did not require citizen scientists to outline valid sections of the sheet. This part of the processing has been automated. We refer to our<a href="https://doi.org/10.5281/zenodo.3378864"> Jungle Weather pre/post-processing repository </a>for more details and example code</p> <p>As such, once digitized and properly aligned the whole record was divided into an estimated 30 million cells and 70 000 header files. Below you find an example of a header file and a table cell. During the Jungle Weather project we selected a subset of ~300K table cells for transcription in efforts to validate further Machine Learning based, automated, transcriptions approaches. All data were transcribed by citizen scientists in the spring/summer of 2020.</p> <p><strong>Notes</strong></p> <p>The provided data is raw data, and expert knowledge is required for the correct interpretation of this data. Please contact the authors for the proper context if you are interested in using this data in your project.</p>
Climatological data from mechanistic model experiments of Boljka and Birner (2022/3; npj Climate and Atmospheric Science)
<p>Some climatological output data from mechanistic dry dynamical core model experiments used for the paper of Boljka and Birner (2022/3): "Potential impact of tropopause sharpness on the structure and strength of the general circulation", npj Climate and Atmospheric Science. For more details see the manuscript. </p>
Supplementary data to `Do science maps from open access literature capture the overall topic structure of an academic field?`
<p>The dataset contains the 8,528 academic articles records related to Sustainable Food research sourced with the query `TS=("sustainab*" NEAR/2 "food*")` .</p> <p>They are the records present in the largest component of the citation network, as specified in the manuscript. </p> <p>The dataset was sourced from OpenAlex based on the original data used in the manuscript and it is composed of the following columns:</p> <table> <tbody> <tr> <td><em><strong>Column</strong></em></td> <td><em><strong>Description</strong></em></td> </tr> <tr> <td>Id</td> <td>OpenAlex ID</td> </tr> <tr> <td>DOI</td> <td>Document Object Identifier</td> </tr> <tr> <td>display_name</td> <td>The article title</td> </tr> <tr> <td>publication_year</td> <td>The publication year of the article</td> </tr> <tr> <td>open_access</td> <td>An object with details of the open access status of the article</td> </tr> </tbody> </table> <p>We choose the `.rdata` format for easy loading in R. Use the function `load()` to add the data frame to the enviroment. </p>
Research data management in the German-speaking Sports Sciences - Survey on the Status Quo
<p>The data set contains survey data on the status quo of research data management within the German-speaking sports science community. The survey was conducted as an online survey in the period from August 16<sup>th</sup> to September 30<sup>th</sup>, 2023.</p>
[DATA_SCIENCE] Interviews PomBase Users, January-February 2016
<p>Here you find the transcripts of interviews collected by Sabina Leonelli as part of the ERC project "The Epistemology of Data-Intensive Science". You also find the information sheet provided to interviewees, which gives you the context for this project. Further information and related publications can be found at www.datastudies.eu. One paper that specifically makes use of these interviews was published by Sabina Leonelli in the journal Philosophy of Science in 2018, under the title "Data in Time: Time-Scales of Data Use in the Life Sciences." The transcripts document yeast researchers' attitudes to data curation and the use of databases in their field. Researchers have consented to have these transcripts made available as Open Data. Other interviewees did not give consent, so those transcripts are held securely by the research team in Exeter.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.