Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
672
datasets available to search
ShareScore release 0.7.1
Dataset results
672 results for “logging”
Wood to Soil 0-10 cm data and Wood to Soil 10-20 cm data to detect the imprint of decaying logs (30-80 cm diameter) from two hurricane cohorts (Hugo, 1989, and Georges, 1998)
Many trees fell during Hurricanes Hugo (1989) and Georges (1998) in Puerto Rico. A debris removal experiment suggested that coarse woody hurricane debris slowed canopy recovery by fueling microbial nitrogen immobilization. We analyzed C, N, microbial biomass C and root length in paired soil samples taken under versus 20-50 cm away from large trunks of two species felled by Hugo and Georges three times during wet and dry seasons during the two years after Georges. Data on soil P and other nutrients have not yet been analyzed. Soil microbial biomass, C and N were higher under than near logs of both age cohorts. Frass from wood boring beetles may induce the early effects. Root length was greater under logs at 0-10 cm depth during the dry season, and away from logs in the wet season, but varied independently of microbial biomass. Thus decaying wood can provide resources exploited by tree roots. Percent soil C and N were significantly higher under than near logs in both the 0-10 and 10-20 cm samples. Microbial biomass C varied significantly among seasons at 0-10 cm depth but differences between positions (under vs away) were only suggestive. Surface soil on the upslope side of the logs had significantly more N and microbial biomass, likely from accumulation of leaf litter above the logs on steep slopes. This study shows that C and N accumulate significantly more in soil under than near decaying logs, even in logs that had only decayed for 7 months, and thus contributes to soil heterogeneity. Tree roots track and exploit resource and nutrient hotspots as they change locations between seasons, so the soil heterogeneity in soil fertility is important for forest productivity. Soil phosphorus (P) availability is most often the most limiting nutrient in wet tropical forests. Total soil P was measured by complete digestion in samples from the upper 10 cm; Olsen extractable P (available) was also measured. Total soil P concentrations were significantly greater under than away fr
Marine Mammal Survey, Sightings and Sampling Event Log at Palmer Station, Antarctica, 2020-2024
Seasonal sea ice-influenced marine ecosystems at both poles are characterized by high productivity concentrated in space and time by local, regional, and remote physical forcing. These polar ecosystems are among the most rapidly changing on Earth. The PALmer (PAL) LTER seeks to build on three decades of long-term research along the western side of the Antarctic Peninsula (WAP) to gain new mechanistic and predictive understanding of ecosystem changes in response to disturbances spanning long-term, subdecadal, and higher-frequency “pulses” driven by a range of processes, including long-term climate warming, natural climate variability, and storms. These disturbances alter food-web composition and ecological interactions across time and space scales that are not well understood. We seek to determine the differential effects of disturbance and resilience on krill predators with different life histories, foraging behaviors, and demographic patterns. Specifically, changes in foraging behavior can affect adult fitness, body condition, and reproductive rates, as well as offspring survival. Preliminary analyses suggest mean chick fledgling mass decreases later in the austral summer as storm disturbances increase. If storms are not a factor influencing chick mass, parental effects or ecosystem phenology may play a larger role. For whales, changes in foraging effort and increases in body condition should correlate with increased pregnancy rates. We will test for linkages between whale foraging efficiency related to storms with female pregnancy rates the following year. Our prediction is that in seasons with more storms and poorer foraging conditions, fewer whales will become pregnant. However, as whales are long-lived, we predict this will not have a major effect on the long-term positive population trend. We will contribute fundamental understanding of how population dynamics and physiological processes are responding within a polar marine ecosystem undergoing profound change
Observation logs for APOGEE-1
<p>These files are the observation logs for the SDSS-III/APOGEE survey. The provide information on when different locations on the sky were observed and for how long. They are to be used with the code at https://github.com/jobovy/apogee for determining the APOGEE selection function (the fraction of potential targets observed as in different parts of the sky).</p>
Circularity3 DDOMP - Project tracksheets: scholarly publications, non-peer-reviewed digital outputs, dataset log, and software log.
<p>Project tracking sheets for the Circularity3 project. </p> <p>The following tracking sheets are provided:</p> <ol> <li>Scholarly Publications </li> <li>Non-peer-reviewed digital outputs</li> <li>Dataset log </li> <li>Software log </li> </ol> <p>For further specifications and explanations of the tracking sheets refer to the Circularity3 DDOMP here: https://doi.org/10.5281/zenodo.11047951</p> <p> </p> <p>This tracking sheets reference widely the PARSEC research teams tracking sheets, to whom we are very grateful for their transparent and insightful documentation.</p> <p>Stall, Shelley, Specht, Alison, Corrêa, Pedro Luiz Pizzigatti, David, Romain, Edmunds, Rorie, Mabile, Laurence, Machicao, Jeaneth, Miyairi, Nobuko, Murayama, Yasuhiro, O'Brien, Margaret, Wyborn, Lesley, & Vellenich, Danton Ferreira. (2023). PARSEC Data and Digital Output Management Plan and Workbook. Zenodo. <a href="https://doi.org/10.5281/zenodo.3891426">https://doi.org/10.5281/zenodo.3891426</a></p>
IODP Expedition 391 Piece log
Dataset includes length data for every whole-round piece: bin length, whole-round piece length (measured by curation staff), and both the archive- and working-half piece lengths (optionally measured by scientists).
IODP Expedition 397T Piece log
Dataset includes length data for every whole-round piece: bin length, whole-round piece length (measured by curation staff), and both the archive- and working-half piece lengths (optionally measured by scientists).
Ship logs from ARCTOS Barents Sea Polar Front 2021-05 cruise
<p>PolarFront 2021-05 ship logs. Original (ISO 8859-1 encoded) text files from the ship logger on Helmer Hanssen.</p>
Antarctic Circumnavigation Expedition sample log: samples collected in the Southern Ocean during the austral summer of 2016/17.
<p><strong>Dataset abstract</strong></p> <p>The Antarctic Circumnavigation Expedition (ACE) spent 90 days circumnavigating Antarctica on the R/V Akademik Tryoshnikov during the austral summer of 2016/17. This dataset provides a record of the samples that were collected during the expedition.</p> <p><strong>Dataset contents</strong></p> <ul> <li>ace_sample_log.csv, data file, comma-separated values</li> <li>README.txt, metadata, text</li> <li>data_file_header.txt, metadata, text</li> </ul> <p><strong>Dataset license</strong></p> <p>This sample log is made available under a Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>
Lago Argentino digital core scans and stratigraphic logs
<p>This dataset includes full resolution (20 micron per pixel) digital core scans of all lake cores collected during the 2019 GCO project coring of Lago Argentino.</p> <p>A second folder includes a stratigraphic log and description of each core, created in PSICAT.</p> <p>This dataset is uploaded alongside the submission "Physical limnology and sediment dynamics of Lago Argentino, the world’s largest ice-contact lake" to JGR: Earth Surface.</p> <p>All analyses were conducted at the Continental Scientific Drilling Facility at the University of Minnesota.</p> <p>For any questions about this dataset, please contact vanwy048@umn.edu .</p>
Relative abundance of the CHC extracts of each population replicate's of I. uriae ticks from Iceland after log centered ratio transformation
<p>Relative abundance of the CHC extracts of each population replicate’s of I. uriae ticks from Iceland after log centered ratio transformation.</p>
2019 search and interaction log from the data catalogue: Research Data Australia
<p>In order to provide a better support to user's data discovery activity, we analysed a data search log in order to understand how data seekers interact with a data search system when they search for data. The data search log is from the research data discovery portal: <a href="https://researchdata.edu.au">Research Data Australia (RDA)</a>. RDA is the data discovery service of the Australian Research Data Commons (ARDC). ARDC is supported by the Australian Government through the National Collaborative Research Infrastructure Strategy Program.</p> <p>Please read the research paper "<a href="https://doi.org/10.1108/JD-12-2021-0245">Large-scale Analysis of Query Logs to Profile Users for Dataset Search</a>" for detailed description and analysis of the datasets, and the software "<a href="https://zenodo.org/record/6321621#.Yh79Tt9xUmA">Python code for processing and clustering a data search log</a>" for the data process and analysis.</p> <p>The search log consists of the entire user-front activity log data for the duration of January to December 2019. During this period, the catalogue contained about 150,000 metadata records of datasets.</p> <p>The dataset (2019_search_log_sessioned.txt) was generated from raw log data with following steps:</p> <ul> <li>Remove entries that were likely from machines instead of human users. Those recorded machine activities may result from downstream aggregators who harvested metadata from RDA by directly sending queries to the catalogue URL instead of using the API endpoint.</li> <li>Identify search sessions from a user - a search session includes all activities a user conducts with a search system in order to satisfy a (information/data) search needs. We followed the following steps to identify search sessions. First, we identified a user by IP address, where a unique IP address was considered a single user. We recognise the limitation of this approach, as several users may share the same IP address, however the IP address is the only information available for identifying a user. <br> Past research in log analysis usually apply the following two methods to identify a session: 30 minutes from the same IP address, and/or more than 30 minutes of inactivity between the current activity event and its immediate preceding event. We examined both methods carefully for our log data and concluded that both ended with large unwanted sessions from machine activities. Therefore, we take a brutal approach, by taking only a session from an IP address with a maximum 30 minutes duration.</li> <li>We also removed sessions whose 40% of activities resulted in ’page not found’ or whose activities were all about accessing grants. Within a session, we removed "duplicated" activities that were exactly as their precedent activity with less than one second time span (this could have been a result of reloading a page).</li> </ul> <p>The dataset (id_to_title_subject.csv) lists title and subject headings per record id.</p> <p> </p>
Geologic time scale - log-spiral
<p>The geologic time scale proportionally represented as a log-spiral. Some key events in Earth's history are marked on the diagram, including major extinction events, global scale glaciations, the intiation of permanent atmospheric oxygen, the formation of the moon, and the formation of Earth's magnetic field. The outer spiral arcs show components of the evolution of life on Earth. There are three versions available each with a light/dark text variant. One version uses the official International Commission on Stratigraphy colour scheme for the time scale and the scientific colours batlow colour scheme for life evolution. The other two versions use scientific colour maps batlow/glasgow and glasgow/batlow for their time scale and history of life evolution respectively. The light variants have black text in transparent regions and should be used for light backgrounds, while the dark variants have white text in transparent zones and should be used on dark backgrounds.</p> <p>Version 1.0.1 corrects a spelling error</p> <p>Version 1.0.2 fixes transposed eon/era labels and adds PDF versions.</p> <p>Version 1.0.3 corrects a spelling error</p> <p>The julia code used to generate the base of the image is provided. Further editing was done in Inkscape.</p>
DFT Calculated xyz and log Files as well as csv Files for Machine Learning in Support of "Tailoring Phosphine Ligands for Improved C H Activation: Insights from Δ-Machine Learning"
<p>Transition metal complexes have played crucial roles in various homogeneous catalytic processes due to their exceptional versatility. This adaptability stems not only from the central metal ions but also from the vast array of choices of the ligand spheres, which form an enormously large chemical space. For example, Rh complexes, with a well-designed ligand sphere, are known to be efficient in catalyzing the C-H activation process in alkanes. To investigate the structure-property relation of the Rh complex and identify the optimal ligand that minimizes the calculated reaction energy ΔE of an alkane C-H activation, we have applied a Δ-Machine Learning method trained on various features to study 1,743 pairs of reactants (Rh(PLP)(Cl)(CO)) and intermediates (Rh(PLP)(Cl)(CO)(H)(propyl)). Our findings demonstrate that the models exhibit robust predictive performance when trained on features derived from electron density (R<sup>2 </sup>= 0.816), and SOAPs (R<sup>2 </sup>= 0.819), a set of position-based descriptors. Leveraging the model trained on xTB-SOAPs that only depend on the xTB-equilibrium structures, we propose an efficient and accurate screening procedure to explore the extensive chemical space of bisphosphine ligands. By applying this screening procedure, <a>we identify ten newly selected reactant-intermediate pairs with an average ΔE </a>of 33.2 kJ mol<sup>-1</sup>, remarkably lower than the average ΔE of the original data set of 68.0 kJ mol<sup>-1</sup>. This underscores the efficacy of our screening procedure in pinpointing structures with significantly lower energy levels.</p> <p>_______________________________________________________________________</p> <p>The dataset contains three file types:</p> <p>Version 1.0:</p> <ol> <li>xyz files of the final optimized Rh-phosphine complexes; one set for the starting materials denoted as "molecule-XXXX_4-times" and one set for the intermediates after C-H activation denoted as "molecule-XXXX_6-times"</li> <li>Gaussian16 log files for the optimization process; one set for the starting materials denoted as "molecule-XXXX_4-times" and one set for the intermediates after C-H activation denoted as "molecule-XXXX_6-times"</li> <li>csv files containing the per molecule features used for training the different machine learning models. The name of the csv files indicates which property was predicted and which model was used</li> </ol> <p>New in version 1.1 (other data is unchanged):</p> <ol> <li>Gaussian16 log files for the ten newly identified bisphosphine ligands; one set for the product material denoted as "LXX_6-times-axial" and one set for the transition state for the C-H activation denoted as "LXX_C-H-activation_TS"</li> </ol>
IODP Expedition 360 Piece log
Dataset includes length data for every whole-round piece: bin length, whole-round piece length (measured by curation staff), and both the archive- and working-half piece lengths (optionally measured by scientists).
Student's logs and perceptions of an automated assessment tool in a software engineering MOOC specialization
<p>Our dataset contains students' perceptions and usage of an automated assessment tool (MOOCauto) for obtaining formative feedback in software engineering assignments that are part of a MOOC specialization at Universidad Politécnica de Madrid (Spain), delivered by the MiriadaX platform. The dataset has previously been used in a study to evaluate students' perceptions of the tool and to analyze their usage patterns using Growth Mixture Models <a href="https://www.computer.org/csdl/magazine/so/5555/01/10196480/1P9AhkBLYXK">(López-Pernas et al., 2023)</a>. The code of each of the assignments is available on Github: <a href="https://github.com/ging-moocs">https://github.com/ging-moocs</a>.</p> <p>Our dataset contains two files:</p> <h2>MOOCauto usage logs</h2> <p>The first file is called<strong> moocauto_logs.csv </strong>and it contains 9,108 anonymized logs of students' use of the automated assessment tool in the MOOC specialization assignments. The columns of the dataset are as follows:</p> <ul> <li><strong>MOOCid</strong>: Unique numeric identifier for the MOOC (1-4)</li> <li><strong>MOOC: </strong>Name of the MOOC: Frontend Development, Backend Development, Git & Github, Fullstack Development</li> <li><strong>AssignmentName</strong>: Name of the assignment.</li> <li><strong>AssignmentId</strong>: Unique identifier for each assignment (1-17)</li> <li><strong>user: </strong>Unique identifier of the student (it varies per assignment)</li> <li><strong>timestamp: </strong>Time in which the assessment was performed</li> <li><strong>score</strong>: Score obtained (0-10)</li> </ul> <h2>Students' perceptions of MOOCauto</h2> <p>The second file is called <strong>moocauto_questionnaire.csv</strong> and it contains 213 students' responses to the questionnaire conducted at the end of each MOOC in order to evaluate their opinion of the tool and perception on usefulness, ease of use, and other aspects related to the Technology Acceptance Model (TAM). The questions were as follows:</p> <ul> <li><strong>What is your general opinion of MOOCauto?</strong> (1 Horrible - 5 Excellent)</li> <li><strong>Indicate your level of agreement with the following statements </strong>(1 Strongly disagree - 5 Strongly agree) <ul> <li>MOOCauto has been easy to install</li> <li>MOOCauto has been easy to use</li> <li>The feedback provided by MOOCauto was easy to understand</li> <li>The feedback provided by MOOCauto was useful</li> <li>The feedback provided by MOOCauto helped me improve my assignments</li> <li>The documentation Of MOOCauto was useful</li> <li>MOOCauto has increased my motivation to work on the assignments</li> <li>I prefer the feedback from MOOCauto than from peer assessment</li> <li>I would like to have a bot like MOOCauto in other MOOCs</li> </ul> </li> <li><strong>How useful do you perceive the following features of MOOCauto?</strong> (1 Useless - 5 Very useful) <ul> <li>It works locally on my computer</li> <li>It allows to run the test suite as many times as I want</li> <li>It provides instantaneous feedback every time the test suite is executed</li> <li>It has documentation that explains its use and available options</li> </ul> </li> </ul>
The SPECIAL Policy Log Vocabulary
<p>This documents specifies <em>splog</em>, a vocabulary to log data processing and sharing events that should comply with a given consent provided by a data subject. We also model the consent actions related to consent giving and revocation.</p> <p> </p> <p>See more at: <a href="http://purl.org/specialprivacy/splog">http://purl.org/specialprivacy/splog</a></p>
Sample FITS file with log linear wavelength solution
<p>This file was wavelength calibrated using IRAF and written to a FITS file using a log linear wavelength solution.</p>
Web robot detection - Server logs
<p>This dataset contains server logs from the search engine of the library and information center of the Aristotle University of Thessaloniki in Greece (<a href="http://search.lib.auth.gr/">http://search.lib.auth.gr/</a>). The search engine enables users to check the availability of books and other written works, and search for digitized material and scientific publications. The server logs obtained span an entire month, from March 1st to March 31 2018 and consist of 4,091,155 requests with an average of 131,973 requests per day and a standard deviation of 36,996.7 requests. In total, there are requests from 27,061 unique IP addresses and 3,441 unique user-agent strings. The server logs are in JSON format and they are anonymized by masking the last 6 digits of the IP address and by hashing the last part of the URLs requested (after last /). The dataset also contains the processed form of the server logs as a labelled dataset of log entries grouped into sessions along with their extracted features (simple semantic features). We make this dataset publicly available, the first one in this domain, in order to provide a common ground for testing web robot detection methods, as well as other methods that analyze server logs.<br> <br> </p>
5405 MHz SigMF baseband recording of RCM-2 (Radarsat Constellation) using PlutoPlus SDR and four log periodic array (LPA) antennas
<p>This dataset contains a recording of the <a href="https://www.asc-csa.gc.ca/eng/satellites/radarsat/technical-features/radarsat-comparison.asp">RCM-2</a> (<a href="https://en.wikipedia.org/wiki/RADARSAT_Constellation">Radarsat Constellation</a>) satellite as it passed over Berkeley, California. The recording was made on 2024-08-13 and is about 15 seconds long, containing acquisition of signal pulses and loss of signal at the tail end of the recording. The dataset is stored in SigMF format. The data files are compressed with xz to reduce their size. The IQ sample rate is 20.0 Msps and the center frequency is 5405 MHz.</p> <p>The linear antenna array used to record contained four HT5 antennas, labeled as: "HT5 antenna UWB log-periodic antenna 1300MHz-10GHz". These four antennas were spaced 21 cm apart. The input from these four antennas was combined with a SP-TX-4B splitter/combiner using equal lengths of LMR400 coax, then amplified using an LNA labeled as "RF AMP 04A: TQP3M9037 0.1-6GHz". The LNA was powered via a +5 volt bias-tee. A "Pluto+" or Pluto Plus SDR was used to sample, with SatDump software. </p> <p> </p>
Core log descriptions and sediment grain size data for Hurricane Ian sediment cores collected in Lee County, Florida, USA
<p>These data represent qualitative and quantitative measurements of sediment cores collected from various environments following the landfall of Hurricane Ian. These sediment cores were collected using pound coring techniques up to 2m into the subsurface to characterize the sedimentological signature of storm deposits resulting from Hurricane Ian. More details regarding these measurements and interpretations of storm deposits can be found in the folllowing manuscript:</p> <p>McCormick, W.M., Briggs, T.R., Hauptman, L.H., Wang, P., Morphologic and sedimentological signatures resulting from Hurricane Ian, southwest Florida, USA: Insight into intra-storm bidirectional sediment transport processes (In Review). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.