Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
22,922
datasets available to search
ShareScore release 0.9.0
Dataset results
22,922 results for “collections as data”
Meteorology and soil moisture data collected at multiple frequencies from the Cross-scale Interactions Study (CSIS) Block-11 site automated monitoring stations: Jornada Basin LTER, 2013 - ongoing
This dataset contains summary data collected at the Jornada Basin LTER program's Cross Scale Interactions Study (CSIS) Block-11 site automated weather station and associated soil substation at several temporal scales. Precipitation data are collected at 1-second frequency during rain events, air temperature and wind are summarized every 5-minutes, and all aboveground sensors are summarized at 30-minute, hourly and daily frequencies. Soil moisture is measured at a 30-minute frequency and summarized daily. Observed values include average/maximum/minimum air temperature, relative humidity, and wind speed; average wind direction; soil moisture, temperature and conductivity. Aboveground sensors are measured and calculated based on 1-second scan rate. Soil moisture is measured every 30-minutes near the weather station and approximately 30-meters distance at a nearby substation. Wind speed is measured at 37cm, 75 cm, 150 cm, and 300 cm, wind direction at approximately 3m, and air temperature and relative humidity at approximately 2.5m. Soil sensors are installed at 10, 20 and 30cm depths.
Meteorology and soil moisture data collected at multiple frequencies from the Cross-scale Interactions Study (CSIS) Block-12 site automated monitoring stations: Jornada Basin LTER, 2013 - ongoing
This dataset contains summary data collected at the Jornada Basin LTER program's Cross Scale Interactions Study (CSIS) Block-12 site automated weather station and associated soil substation at several temporal scales. Air temperature and wind are summarized every 5-minutes, and all aboveground sensors are summarized at 30-minute, hourly and daily frequencies. Soil moisture is measured at a 30-minute frequency and summarized daily. Observed values include average/maximum/minimum air temperature, relative humidity, and wind speed; average wind direction; soil moisture, temperature and conductivity. Aboveground sensors are measured and calculated based on 1-second scan rate. Soil moisture is measured every 30-minutes near the weather station and approximately 30-meters distance at a nearby substation. Wind speed is measured at 37cm, 75 cm, 150 cm, and 300 cm, wind direction at approximately 3m, and air temperature and relative humidity at approximately 2.5m. Soil sensors are installed at 10, 20 and 30cm depths. The nearest precipitation data is available from CSIS Block-13 site automated weather station located 290m distance.
Meteorology and soil moisture data collected at multiple frequencies from the Cross-scale Interactions Study (CSIS) Block-13 site automated monitoring stations: Jornada Basin LTER, 2013 - ongoing
This dataset contains summary data collected at the Jornada Basin LTER program's Cross Scale Interactions Study (CSIS) Block-13 site automated weather station and associated soil substation at several temporal scales. Precipitation data are collected at 1-second frequency during rain events, air temperature and wind are summarized every 5-minutes, and all aboveground sensors are summarized at 30-minute, hourly and daily frequencies. Soil moisture is measured at a 30-minute frequency and summarized daily. Observed values include average/maximum/minimum air temperature, relative humidity, and wind speed; average wind direction; soil moisture, temperature and conductivity. Aboveground sensors are measured and calculated based on 1-second scan rate. Soil moisture is measured every 30-minutes near the weather station and approximately 30-meters distance at a nearby substation. Wind speed is measured at 37cm, 75 cm, 150 cm, and 300 cm, wind direction at approximately 3m, and air temperature and relative humidity at approximately 2.5m. Soil sensors are installed at 10, 20 and 30cm depths.
Meteorology and soil moisture data collected at multiple frequencies from the Cross-scale Interactions Study (CSIS) Block-14 site automated monitoring stations: Jornada Basin LTER, 2013 - ongoing
This dataset contains summary data collected at the Jornada Basin LTER program's Cross Scale Interactions Study (CSIS) Block-14 site automated weather station and associated soil substation at several temporal scales. Precipitation data are collected at 1-second frequency during rain events, air temperature and wind are summarized every 5-minutes, and all aboveground sensors are summarized at 30-minute, hourly and daily frequencies. Soil moisture is measured at a 30-minute frequency and summarized daily. Observed values include average/maximum/minimum air temperature, relative humidity, and wind speed; average wind direction; soil moisture, temperature and conductivity. Aboveground sensors are measured and calculated based on 1-second scan rate. Soil moisture is measured every 30-minutes near the weather station and approximately 30-meters distance at a nearby substation. Wind speed is measured at 37cm, 75 cm, 150 cm, and 300 cm, wind direction at approximately 3m, and air temperature and relative humidity at approximately 2.5m. Soil sensors are installed at 10, 20 and 30cm depths.
Meteorology and soil moisture data collected at multiple frequencies from the Cross-scale Interactions Study (CSIS) Block-15 site automated monitoring stations: Jornada Basin LTER, 2013 - ongoing
This dataset contains summary data collected at the Jornada Basin LTER program's Cross Scale Interactions Study (CSIS) Block-15 site automated weather station and associated soil substation at several temporal scales. Precipitation data are collected at 1-second frequency during rain events, air temperature and wind are summarized every 5-minutes, and all aboveground sensors are summarized at 30-minute, hourly and daily frequencies. Soil moisture is measured at a 30-minute frequency and summarized daily. Observed values include average/maximum/minimum air temperature, relative humidity, and wind speed; average wind direction; soil moisture, temperature and conductivity. Aboveground sensors are measured and calculated based on 1-second scan rate. Soil moisture is measured every 30-minutes near the weather station and approximately 30-meters distance at a nearby substation. Wind speed is measured at 37cm, 75 cm, 150 cm, and 300 cm, wind direction at approximately 3m, and air temperature and relative humidity at approximately 2.5m. Soil sensors are installed at 10, 20 and 30cm depths.
Heliconia collection data
Mineral concentrations of Heliconia bract fluid and organic matter accumulated in the bracts. Data collected during a study of the invertebrates in Heliconia bract fluid. Support for this work was provided by grants BSR-8811902, DEB-9411973, DEB-9705814 , DEB-0080538, DEB-0218039 , DEB-0620910 , DEB-1239764, DEB-1546686, and DEB-1831952 from the National Science Foundation to the University of Puerto Rico as part of the Luquillo Long-Term Ecological Research Program. Additional support provided by the University of Puerto Rico and the International Institute of Tropical Forestry, USDA Forest Service.
Limno run codes, dates, and locations associated with vertical lake profile data collected in the McMurdo Dry Valleys, Antarctica (1991-2025, ongoing)
This data package provides a summary of "limno runs" performed each season in lakes located throughout the McMurdo Dry Valleys of Antarctica as part of the McMurdo Dry Valleys Long Term Ecological Research program. Each limno run is identified by unique code that allows one to compile a complete set of limnological data from a particular lake and location at specific time points. There are usually two or three limno runs performed per lake per austral summer field season, although the number of runs may vary by lake and by season depending on site access, site conditions, and other external factors.
Skin-blubber biopsy samples and associated demographic data collected from cetaceans encountered along the Western Antarctic Peninsula (WAP), 2010 – 2024
Baleen whale populations in the Southern Ocean are recovering after intense commercial whaling in the 20th century. Along the Western Antarctic Peninsula (WAP), this recovery is occurring in one of the planet's most rapidly changing marine ecosystems. Understanding how climate-driven changes influence the population dynamics of whales in this region is critical for understanding what conservation and management actions must be prioritized to maintain the structure and function of this marine ecosystem. To begin understanding the dynamics of whale recovery under continued environmental change, we need to study these whales' demography and population dynamics. As part of our annual sampling surveys for cetaceans along the WAP through the PAL LTER program, we actively collect remote non-lethal skin-blubber biopsy samples and have developed one of the most extensive tissue archives in the Southern Ocean. With these samples, we conduct a series of demographic and physiological measurements. Using the skin portion of the biopsy sample, we isolate nuclear and mitochondrial DNA (mtDNA) to develop a DNA profile for each sample, including genetic sex, a microsatellite genotype, and a mtDNA haplotype. These profiles are used to compare sex ratios of the population, determine individual recaptures through genotype analysis, and better understand population mixing. Using the blubber portion of the biopsy sample, we isolate endocrine markers (e.g., progesterone and cortisol) to monitor population pregnancy rates and stress levels. This data represents some of the first non-lethal quantitative observations of the demography and population dynamics of recovering whale populations in the Antarctic and provides a critical reference point for future work as the Antarctic climate continues to change and populations continue to recover from whaling.
Temperature and ADCP data collected on Lake Geneva between 2015 and 2017
<p>Data collected between 2015 and 2017 on Lake Geneva by Acoustic Doppler Current Profiler (ADCP) and CTDs. One file includes all the temperature profiles, the two others are the ADCP data (up- and down-looking) at the SHL2 station (centre of the main basin). Coordinates of the SHL2 station are 534700 and 144950 in the Swiss CH1903 coordinate system. The file with the CTD data contains the coordinates of the sample location (lat, lon), times (in MATLAB time), depths (in meters) and temperatures (in °C).</p> <p>All files are in MATLAB .mat format.</p>
JUMP - Data collection - Part II: Zonal jets using three different approaches, laboratory - Global Climate Models - observations.
<p>The formation of large scale structures in three-dimensional (3D) turbulent flows. How small-scale dynamics organize in turbulent flows to grow large scale coherent circulation? is at the heart of fundamental studies in fluid dynamics. It appears to be equally important for our understanding of atmospheric dynamics, oceanography, meteorology and more generally geophysical fluid dynamics. Here, we deliver a data collection that <strong>(1)</strong> gathers measurements of 3D turbulent flows that emulate planetary atmospheres of the gas giants. Turbulent flows are explored using three different approaches, laboratory experiments, numerical simulations and direct planetary observations. All data set are computed in order to easily extract flow properties, i.e. high resolution maps of the different velocity components and flow vorticity (useful for further diagnostic). The data collected are fully discribed in Cabanes et al GRL (2020) "Revealing the intensity of turbulent energy transfer in planetary atmospheres" and can be used to compute <strong>(2)</strong> theoretical diagnostics with the numerical codes that allow to reveal the physical meaning of flow measurements. Numerical codes are available on https://github.com/scabanes</p> <p>We deliver (1) data collection and (2) numerical codes in the following files attached:</p> <p>(1) Data collection:</p> <ul> <li>A PDF file named <strong>JUMP-zonal-jets-data-collection-GRL.pdf</strong> that describes the following data files and nomenclature.</li> <li>A zip File of the velocity fields in the lab, interpolated on Polar and Cartesian grids <ul> <li><strong>JUMP-JetsInTheLab.zip</strong></li> </ul> </li> <li>A netcdf file of velocity fields of our Saturn reference simulation <ul> <li><strong>uvData-SRS-istep-312000-nstep-50-niz-12.nc</strong></li> </ul> </li> <li>Two netcdf files of velocity fields from Cassini observations of Jupiter<strong> </strong> <ul> <li><strong>uvData-JupObs-istep-0-nstep-4-niz-1.nc</strong></li> <li><strong>StatisticalData-JupObs.nc</strong></li> </ul> </li> <li>A zip file of potential vorticity profiles for Saturn and Jupiter observations <ul> <li><strong>IPV-QGPV-Jupiter-Saturn.zip</strong></li> </ul> </li> </ul> <p>(2) Numerical codes:</p> <ul> <li>Codes for statistical analysis in spherical geometry on Github. --> <a href="https://www.google.com/url?q=https%3A%2F%2Fgithub.com%2Fscabanes%2FPOST&sa=D&sntz=1&usg=AFQjCNFuDU0eij4XGxQfReO92CHfJz6PBA">https://github.com/scabanes/POST</a></li> <li>Codes for statistical analysis in cylindrical geometry on Github. --> <a href="https://www.google.com/url?q=https%3A%2F%2Fgithub.com%2Fscabanes%2FJUMP&sa=D&sntz=1&usg=AFQjCNGUQ1YIFhSxBAg4Hl_5gOLB_4LxLA">https://github.com/scabanes/JUMP</a></li> <li>Codes for statistical analysis in cartesian geometry on Github. --> <a href="https://www.google.com/url?q=https%3A%2F%2Fgithub.com%2Fscabanes%2FJUMP&sa=D&sntz=1&usg=AFQjCNGUQ1YIFhSxBAg4Hl_5gOLB_4LxLA">https://github.com/scabanes/JUMP</a></li> </ul> <p> </p> <p>The purpose of this data collection is to reveal statistical properties of planetary flows. By computing the same analysis on different data sets the researcher allows direct confrontation of planetary observations with idealized laboratory and numerical models. Idealized models are specially designed to sweep on a large array of parameters in order to understand what parameters control planetary global circulation. The data collected and generated by the researcher deliver <strong>(1)</strong> velocity measurements of 3D turbulent flows using the different approaches (observations-laboratory-numerics) and <strong>(2)</strong> guidelines to compute the appropriate statistical analysis through the PTST. Here, the ground-breaking novelty is that the researcher deliver the possibility to compute statistical diagnostics adapted to the different geometries: the spherical geometry of planetary flows, i.e. 2D latitude-longitude maps, the cylindrical geometry of laboratory experiments, i.e. 2D flows in a rotating cylindrical tank, and the Cartesian geometry of idealized numerical simulations. Indeed, the math behind each statistical diagnostics must account for the different geometrical configurations in order to properly confront the different approaches. The PTST is also designed to be easily re-used by different communities such as experimentalists, numericists and atmosphericists that deal with 3D or 2D turbulent flows.</p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>This project has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement N° 797012.</p>
1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test using artificial intelligence and image-based pattern recognition techniques
<p><strong>The Great Isaiah Scroll (1QIsa<sup>a</sup>) data set for writer identification</strong></p> <p>This data set is collected for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497</p> <p>Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> <br> <strong>Copyright (c) </strong> University of Groningen, 2021. All rights reserved.<br> <strong>Disclaimer and copyright notice for all data contained on this .tar.gz file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the data for research purposes. It is not allowed to distribute this data for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this data.</p> <p><strong>4) </strong>the user should refer to the first public article on this data set:<br> <br> <em>Popović, M., Dhali, M. A., & Schomaker, L. (2020). Artificial intelligence-based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.</em><br> <br> BibTeX:</p> <pre>@article{popovic2020artificial, title={Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)}, author={Popovi{\'c}, Mladen and Dhali, Maruf A and Schomaker, Lambert}, journal={arXiv preprint arXiv:2010.14476}, year={2020} }</pre> <p><strong>5) </strong>the recipient should refrain from proliferating the data set to third parties external to his/her local research group. Please refer interested researchers to this site for obtaining their own copy.</p> <p><strong>Organisation of the data:</strong></p> <p>The .tar.gz file contains three directories: images, features, and plots. The included 'README' file contains all the instructions.</p> <p>The 'images' directory contains NetPBM images of the columns of 1QIsa<sup>a</sup>. The NetPBM format is chosen because of its simplicity. Additionally, there is no doubt about lossy compression in the processing chain. There are two images for each of the Great Isaiah Scroll columns: one is the direct binarized output from the BiNet (<em>arxiv.org/abs/1911.07930</em>) system, and the other one is the manually cleaned version of the binarized output. The file names for the direct binarized output are of the format '1QIsaa_col<columnnr>.pbm', for example, '1QIsaa_col15.pbm'. And, for the cleaned version, the format is '1QIsaa_col<columnnr>_cleaned.pbm', for example, '1QIsaa_col15_cleaned.pbm'. Note: the image files are not in a separate directory; they will be extracted in the same place. However, due to the unique naming, there is no problem extracting them in one single directory.</p> <p>The 'features' directory contains feature files computed for each of the column images. There are two types of feature files: Hinge and Adjoined. They are distinguishable by their extension, for example, '1QIsaa_col15_cleaned.hinge' and '1QIsaa_col15_cleaned.adjoined'. They are also arranged in separate directories for ease of use.</p> <p>The 'plots' directory contains a simple python script to perform PCA on the feature files and then visualize them in a 3D plot. The file takes the location of feature files as an input. The 'README_plot' file contains examples of how-to-run in the terminal.</p> <p><strong>Brief description:</strong><br> According to ImageMagick's' identify' tool, the original images are in grayscale (.jpg) from Brill collection, in '8-bit Gray 256c'. These images pass through multiple preprocessing measures to become suitable for pattern recognition-based techniques. The first step in preprocessing is the image-binarization technique. In order to prevent any classification of the text-column images based on irrelevant background patterns, a specific binarization technique (BiNet) was applied, keeping the original ink traces intact. After performing the binarization, the images were cleaned further by removing the adjacent columns that partially appear on the target columns' images. Finally, few minor affine transformations and stretching corrections were performed in a restrictive manner. These corrections are also targeted for aligning the texts where the text lines get twisted due to the leather writing surface's degradation. Hence, the clean images are there in the directory along with the direct binarized images. No effort has been made to obtain a balanced set in any way.</p> <p><strong>Tools:</strong><br> <strong>Binarization:</strong><br> The BiNet tool is available for scientific use upon request (m.a.dhal(at)rug.nl)</p> <p><strong>Image Morphing:</strong><br> In the original article, data augmentation was performed using image morphing. The tool is available on GitHub:<br> https://github.com/GrHound/imagemorph.c</p> <p><strong>Features for writer identification:</strong><br> Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> <em><strong>1. </strong>L. Schomaker & M. Bulacu (2004). Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.<br> <strong>2. </strong>Bulacu, M. & Schomaker, L.R.B. (2007). Text-independent Writer Identification and Verification Using Textural and Allographic Features, IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</em><br> <br> The features (hinge, fraglets) have been combined in a single MS Windows application, GIWIS, which is available for scientific use upon request (l.r.b.schomaker(at)rug.nl)</p> <p><strong>If you have any question, please contact us:</strong><br> Maruf A. Dhali <m.a.dhali(at)rug.nl><br> Lambert Schomaker <l.r.b.schomaker(at)rug.nl><br> Mladen Popović <m.popovic(at)rug.nl></p> <p><strong>Please cite our papers if you use this data set:</strong><br> <em><strong>1.</strong> Popović, M., Dhali, M. A., & Schomaker, L. (2020). Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.<br> <strong>2. </strong>Dhali, M. A., de Wit, J. W., & Schomaker, L. (2019). Binet: Degraded-manuscript binarization in diverse document textures and layouts using deep encoder-decoder networks. arXiv preprint arXiv:1911.07930.</em></p>
Conductivity–Temperature–Depth (CTD) and dissolved oxygen profile data from shipboard surveys collected within Olympic Coast National Marine Sanctuary, 2005-2023
<p>This data set includes Conductivity-Temperature-Depth (CTD) and dissolved oxygen profile data that were collected along Washington State’s outer coast within Olympic Coast National Marine Sanctuary towards the northernmost extent of the California Current System. Measurements were made at fourteen hydrographic stations during mooring deployment, recovery, and maintenance cruises between the months of May and October from 2005–2023. The 792 CTD profiles were acquired using Sea-Bird Scientific 19 SeaCAT or 19plus SeaCAT CTD profilers with associated SBE-43 (Sea-Bird Electronics) or Beckman or YSI-type (Yellow Springs Instruments) dissolved oxygen sensors. The data were processed via Sea-Bird Scientific’s SBE Data Processing application using six of the modules in the following order: <em>Data Conversion, Filter, Align CTD, Loop Edit, Derive, and Bin Average</em>. These processing steps and associated methods are the same as those used to process CTD data that make up the <a href="../records/5814071">Newport Hydrographic Line time series</a> located off the central Oregon coast thus allowing for a direct comparison between the two regions.</p> <table> <tbody> <tr> <td><strong>Station Name </strong></td> <td><strong>Latitude</strong></td> <td><strong>Longitude</strong></td> <td><strong>Water Depth (m, MLLW)</strong></td> </tr> <tr> <td><strong>Makah Bay (MB)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>MB015</td> <td>48.3254oN</td> <td>124.6768oW</td> <td>15</td> </tr> <tr> <td>MB042</td> <td>48.3240oN</td> <td>124.7354oW</td> <td>42</td> </tr> <tr> <td><strong>Cape Alava (CA)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>CA015</td> <td>48.1663oN</td> <td>124.7568oW</td> <td>15</td> </tr> <tr> <td>CA042</td> <td>48.1660oN</td> <td>124.8234oW</td> <td>42</td> </tr> <tr> <td>CA065 </td> <td>48.1659oN</td> <td>124.8949oW</td> <td>65</td> </tr> <tr> <td><strong>Teahwhit Head (TH)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>TH015</td> <td>47.8761oN</td> <td>124.6195oW</td> <td>15</td> </tr> <tr> <td>TH042</td> <td>47.8762oN</td> <td>124.7334oW</td> <td>42</td> </tr> <tr> <td>TH065 </td> <td>47.8767oN</td> <td>124.7967oW</td> <td>65</td> </tr> <tr> <td><strong>Kalaloch (KL)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>KL015</td> <td>47.6008oN</td> <td>124.4284oW</td> <td>15</td> </tr> <tr> <td>KL027</td> <td>47.5946oN</td> <td>124.4971oW</td> <td>27</td> </tr> <tr> <td>KL050 </td> <td>47.5933oN</td> <td>124.6112oW</td> <td>50</td> </tr> <tr> <td><strong>Cape Elizabeth (CE)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>CE015</td> <td>47.3568oN</td> <td>124.3481oW</td> <td>15</td> </tr> <tr> <td>CE042</td> <td>47.3531oN</td> <td>124.4887oW</td> <td>42</td> </tr> <tr> <td>CE065 </td> <td> 47.3528oN</td> <td>124.5669oW</td> <td>65</td> </tr> </tbody> </table>
RAYUELA - Open Data - Data collected through a serious game created to identify patterns and profiles of young potential victims/perpetrators of cybercrimes.
<p>The data of this dataset have been collected in the pilots carried out by the RAYUELA project in different countries of the European Union. The participants are minors and the game sessions have been carried out in schools and summer camps in a supervised way.</p>
UC Santa Barbara Invertebrate Zoology Collection (UCSB-IZC) Data Archive and Biodiversity Dataset Graph hash://md5/10663911550bb52a0f5741993f82db9d hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c
<p>A biodiversity dataset graph: UCSB-IZC</p> <p>The intended use of this archive is to facilitate (meta-)analysis of the UC Santa Barbara Invertebrate Zoology Collection (UCSB-IZC). UCSB-IZC is a natural history collection of invertebrate zoology at Cheadle Center of Biodiversity and Ecological Restoration, University of California Santa Barbara.</p> <p>This dataset provides versioned snapshots of the UCSB-IZC network as tracked by Preston [2,3] between 2021-10-08 and 2021-11-04 using [preston track "https://api.gbif.org/v1/occurrence/search/?datasetKey=d6097f75-f99e-4c2a-b8a5-b0fc213ecbd0"].</p> <p>This archive contains 14349 images related to 32533 occurrence/specimen records. See included sample-image.jpg and their associated meta-data sample-image.json [4].</p> <p>The images were counted using:</p> <p>$ preston cat hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c\<br> | grep -o -P ".*depict"\<br> | sort\<br> | uniq\<br> | wc -l</p> <p>And the occurrences were counted using:</p> <p>$ preston cat hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c\<br> | grep -o -P "occurrence/([0-9])+"\<br> | sort\<br> | uniq\<br> | wc -l</p> <p>The archive consists of 256 individual parts (e.g., preston-00.tar.gz, preston-01.tar.gz, ...) to allow for parallel file downloads. The archive contains three types of files: index files, provenance files and data files. Only two index and provenance files are included and have been individually included in this dataset publication. Index files provide a way to links provenance files in time to establish a versioning mechanism.</p> <p>To retrieve and verify the downloaded UCSB-IZC biodiversity dataset graph, first download preston-*.tar.gz. Then, extract the archives into a "data" folder. Alternatively, you can use the Preston [2,3] command-line tool to "clone" this dataset using:</p> <p>$ java -jar preston.jar clone --remote https://archive.org/download/preston-ucsb-izc/data.zip/,https://zenodo.org/record/5557670/files,https://zenodo.org/record/5660088/files/</p> <p>After that, verify the index of the archive by reproducing the following provenance log history:</p> <p>$ java -jar preston.jar history<br> <urn:uuid:0659a54f-b713-4f86-a917-5be166a14110> <http://purl.org/pav/hasVersion> <hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36> .<br> <hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c> <http://purl.org/pav/previousVersion> <hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36> .</p> <p>To check the integrity of the extracted archive, confirm that each line produce by the command "preston verify" produces lines as shown below, with each line including "CONTENT_PRESENT_VALID_HASH". Depending on hardware capacity, this may take a while.</p> <p>$ java -jar preston.jar verify<br> hash://sha256/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c file:/home/jhpoelen/ucsb-izc/data/ce/1d/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c OK CONTENT_PRESENT_VALID_HASH 66438 hash://sha256/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c<br> hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844 file:/home/jhpoelen/ucsb-izc/data/f6/8d/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844 OK CONTENT_PRESENT_VALID_HASH 4093 hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844<br> hash://sha256/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef file:/home/jhpoelen/ucsb-izc/data/3e/70/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef OK CONTENT_PRESENT_VALID_HASH 5746 hash://sha256/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef<br> hash://sha256/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b file:/home/jhpoelen/ucsb-izc/data/99/58/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b OK CONTENT_PRESENT_VALID_HASH 6147 hash://sha256/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b</p> <p>Note that a copy of the java program "preston", preston.jar, is included in this publication. The program runs on java 8+ virtual machine using "java -jar preston.jar", or in short "preston".</p> <p>Files in this data publication:</p> <p>--- start of file descriptions ---</p> <p>-- description of archive and its contents (this file) --<br> README</p> <p>-- executable java jar containing preston [2,3] v0.3.1. --<br> preston.jar</p> <p>-- preston archive containing UCSB-IZC (meta-)data/image files, associated provenance logs and a provenance index --<br> preston-[00-ff].tar.gz</p> <p>-- individual provenance index files --<br> 2a5de79372318317a382ea9a2cef069780b852b01210ef59e06b640a3539cb5a</p> <p>-- example image and meta-data --<br> sample-image.jpg (with hash://sha256/916ba5dc6ad37a3c16634e1a0e3d2a09969f2527bb207220e3dbdbcf4d6b810c)<br> sample-image.json (with hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844)</p> <p>--- end of file descriptions ---</p> <p><br> References</p> <p>[1] Cheadle Center for Biodiversity and Ecological Restoration (2021). University of California Santa Barbara Invertebrate Zoology Collection. Occurrence dataset https://doi.org/10.15468/w6hvhv accessed via GBIF.org on 2021-11-04 as indexed by the Global Biodiversity Informatics Facility (GBIF) with provenance hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36 hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c.<br> [2] https://preston.guoda.bio, https://doi.org/10.5281/zenodo.1410543 .<br> [3] MJ Elliott, JH Poelen, JAB Fortes (2020). Toward Reliable Biodiversity Dataset References. Ecological Informatics. https://doi.org/10.1016/j.ecoinf.2020.101132<br> [4] Cheadle Center for Biodiversity and Ecological Restoration (2021). University of California Santa Barbara Invertebrate Zoology Collection. Occurrence dataset https://doi.org/10.15468/w6hvhv accessed via GBIF.org on 2021-10-08. https://www.gbif.org/occurrence/3323647301 . hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844 hash://sha256/916ba5dc6ad37a3c16634e1a0e3d2a09969f2527bb207220e3dbdbcf4d6b810c</p>
Fatiando a Terra data v1.0.0: A curated collection of open geophysics data for tutorials and documentation
<p>This repository holds curated sample datasets that can be used in the documentation and tutorials of the <a href="https://www.fatiando.org/">Fatiando a Terra</a> project. All datasets are cleaned and formatted versions of openly available data under permissive licenses or in the public domain.</p> <p>More information about datasets and the code for cleaning, formatting, and preprocessing the data can be found at: <a href="https://github.com/fatiando/data">https://github.com/fatiando/data</a></p> <p>See the README.md file for information on data sources and their original licenses.</p> <p><strong>NOTE:</strong> This collection uses <a href="https://semver.org/">semantic versioning</a> (i.e., MAJOR.MINOR.BUGFIX). Major releases mean that backwards incompatible changes were made to the data. Minor releases add new data without changing existing files. Bug fix releases fix errors in a previous release that makes the data unusable. Changes to the current data files will always be published as a major release unless the file(s) in the previous release was unusable/corrupted.</p>
Spatially gridded cross-shelf hydrographic sections and monthly climatologies from shipboard survey data collected along the Newport Hydrographic Line, 1997-2021
<p>This data set, described in detail in <a href="https://www.sciencedirect.com/science/article/pii/S2352340922001342">Risien et al. (2022)</a>, contains Newport Hydrographic Line station data; gridded, cross-shelf hydrographic sections; and derived monthly climatologies for temperature, practical salinity, potential density, spiciness, and dissolved oxygen. It consists of CSV (Comma Separated Values) files (<em>newport_hydrographic_line_station_data</em><em>.</em><em>zip</em>) that contain CTD observations collected at the seven hydrographic stations located 1, 3, 5, 10, 15, 20 and 25 nautical miles west of Newport, Oregon between March 1997 and July 2021. Additionally, the data set contains three NetCDF files that follow CF (Climate and Forecast) metadata conventions: <em>newport_hydrographic_line_gridded_sections</em><em>.nc</em> contains observations gridded to a 0.01<sup>o</sup> x 1 dbar longitude - pressure grid to create cross-shelf hydrographic sections for each of the five variables for each cruise. <em>newport_hydrographic_line_gridded_section_climatologies</em><em>.nc</em> contains climatological hydrographic sections, calculated using harmonic analysis over the 24-year period March 1997 to February 2021 and reported here for the middle of each month, and <em>newport_hydrographic_line_gridded_section_coefficients.nc</em> contains the associated linear regression model coefficients for all five variables. From the regression coefficients, users can construct seasonal cycles at any location in the gridded section with a temporal resolution that best suits their specific needs. Finally, this data set includes example MATLAB and R scripts that show how to read the data files, plot cross-shelf hydrographic sections, and calculate daily and monthly climatologies using the regression coefficients.</p>
Single-crystal X-ray diffractometry data for a sample of [Cu(HF₂)(pyrazine)₂]PF₆ collected on beamline I19-2 at Diamond Light Source
<p>Single-crystal X-ray diffractometry data for a sample of [Cu(HF₂)(pyrazine)₂]PF₆.</p> <p>These data were collected at Diamond Light Source, on beamline I19 (experiments hutch 2), on 2022-01-30, and are particularly useful for testing data reduction routines. They are known to produce good merging statistics and final structure refinement.</p> <p>The sample was prepared as follows:<br> Ammonium hexafluorophosphate (NH₄PF₆) (0.310 g, 1.9 mmol), ammonium hydrogen difluoride ((NH₄)HF₂) (0.109 g, 1.9 mmol) and pyrazine (C₄H₄N₂) (0.300 g, 3.7 mmol) were dissolved in 5 mL of deionised water. The obtained colourless solution was slowly added to a blue solution of copper(II) nitrate prepared by dissolving copper(II) nitrate hemipentahydrate (Cu(NO₃)₂ · 2.5(H₂O)) (0.425 g, 1.8 mmol) in 5 mL of deionised water. The solutions were mixed in a plastic beaker at room temperature. The formation of blue crystals of [Cu(HF₂)(pyrazine)₂]PF₆ on the side of the beaker started after few seconds and continued for about 24 hours during which the sealed beaker was not moved.</p> <p>The sample was measured at room temperature and the illuminating beam had a wavelength of 0.4859 Å (25.52 keV).</p> <p>Beamline I19-2 at Diamond Light Source, a four-circle κ-geometry diffractometer (see <a href="https://onlinelibrary.wiley.com/doi/10.1107/97809553602060000936">[Kern 2019]</a>) with an undulator source, is described in <a href="https://doi.org/10.1107/S0909049512008801">[Nowell 2012]</a> but has since been upgraded to use a Dectris Eiger2 X 4M CdTe hybrid photon counting detector. The data are written in the <a href="https://manual.nexusformat.org/classes/applications/NXmx.html">NXmx variant</a> of the <a href="https://www.nexusformat.org/">NeXus format</a>, and so include metadata with a functionally complete description of the diffractometer.</p> <p>Inventory of data:</p> <ul> <li><strong><code>01_CuHF2pyz2PF6b_Phi.tar.xz</code></strong><br> A single 1750-image 350° φ rotation scan from -175° to 175° with 0.2° rotation per image, an exposure time of 0.1 s per image, ω = -90°, κ = 0° and 2θ = 0°.</li> <li><strong><code>02_CuHF2pyz2PF6b_2T.tar.xz</code></strong><br> A single 1750-image 350° φ rotation scan from -175° to 175° with 0.2° rotation per image, an exposure time of 0.1 s per image, ω = -90°, κ = 0° and 2θ = 20°.</li> <li><strong><code>03_CuHF2pyz2PF6b_P_O.tar.xz</code></strong><br> Two sequential rotation scans: <ul> <li><strong><code>CuHF2pyz2PF6b_P_O_01.nxs</code></strong><br> A 1750-image 350° φ scan from -175° to 175° with ω = -90°, κ = 0° and 2θ = 0°.</li> <li><strong><code>CuHF2pyz2PF6b_P_O_02.nxs</code></strong><br> A 600-image 120° ω scan from -125° to -5° with φ = -90°, κ = 45° and 2θ = 0°.</li> </ul> Both scans had 0.2° rotation per image and an exposure time of 0.1 s per image.</li> </ul> <p>The same sample was used for all these measurements. Throughout, the sample-to-detector distance was 85 mm and the beam was attenuated to 0.2% of its full intensity.</p> <p>For each rotation scan, the data comprise a single top-level NXmx-format NeXus file named <code><filename>.nxs</code>, one or more image files named <code><filename>_00000n.h5</code>, where <code>n</code> is a numeral, and a single detector metadata file named <code><filename>_meta.h5</code>. The NeXus file contains an HDF5 virtual data set that links to the data in the image file(s), and several HDF5 external links to data in the detector metadata file.</p> <p>For internal reference of Diamond Light Source staff, these data were collected as part of commissioning visit CM31144-1. Some file names and corresponding HDF5 link targets have been altered from their original names for consistency with the file contents.</p>
Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials
<p>Toxicogenomics (TGx) approaches are increasingly applied to gain insight into the possible toxicity mechanisms of engineered nanomaterials (ENMs). Omics data can be valuable to elucidate the mechanism of action of chemicals and develop predictive models in toxicology. While vast amounts of transcriptomics data from ENM exposures have already been accumulated, a unified, easily accessible and reusable collection of transcriptomics data for ENMs is currently lacking. In an attempt to improve the FAIRness of already existing transcriptomics data for nanomaterials, we curated a collection of homogenized transcriptomics data from human, mouse and rat ENM exposures <em>in vitro</em> and <em>in vivo</em>.</p>
Sea ice core temperature and salinity data collected during the 2019 SCALE Winter Cruise
<p>Temperature and salinity profiles of sea ice cores extracted from in situ sea ice floes and lifted pancakes were measured in the Atlantic sector of the Antarctic Marginal Ice Zone during the Southern oCean seAsonal Experiment (SCALE) winter cruise in 2019 (<a href="http://www.scale.org.za">www.scale.org.za</a>) aboard the SA Agulhas II.</p>
Sea ice core temperature and salinity data collected during the 2019 SCALE Spring Cruise
<p>Temperature and salinity profiles of sea ice cores extracted from in situ sea ice floes and lifted pancakes were measured in the Atlantic sector of the Antarctic Marginal Ice Zone during the Southern oCean seAsonal Experiment (SCALE) spring cruise in 2019 (<a href="http://www.scale.org.za">www.scale.org.za</a>) aboard the SA Agulhas II.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.