Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,481 results for “data set”

Learn how ShareScore rates datasets ↗
zenodo44/100

Coordinates and checklists of alien species populations as obtained from the DASCO workflow and the SInAS data set

<p>This data set contains coordinate records of alien (i.e., non-native) species populations worldwide and aggregated checklists of alien species for individual regions. The regions consists of non-overlapping polygons representing countries, sub-national or coastal marine ecoregions.&nbsp;</p><p>The data set was produced by applying the DASCO workflow (https://doi.org/10.5281/zenodo.5841930) using the SInAS database (version 2.5; https://doi.org/10.5281/zenodo.10038256). The workflow imports checklists of alien species such as those stored in SInAS, and extracts coordinates for the alien regions (according to SInAS) from GBIF and OBIS. After cleaning and thinning the coordinates, the workflow exports a list of coordinates of alien populations for all species included in SInAS and with records on GBIF or OBIS.</p><p>These files are part of a manuscript published in the journal Neobiota, where the workflow is described in detail (Seebens &amp; Kaplan 2022, https://doi.org/10.3897/neobiota.74.81082).</p><p>DASCO_AlienCoordinates_SInAS_2.5.gz contains the coordinates of alien populations.</p><p>DASCO_AlienRegions_SInAS_2.5.csv contains the checklists of alien species per region. Note that this only includes species with GBIF and OBIS records. For more comprehensive checklists, other databases such as those listed here (https://doi.org/10.5281/zenodo.10038256) should be consulted.</p><p>OBIS_SpeciesKeys_SInAS_2.5.csv contains the species keys from OBIS.</p><p>GBIF_SpeciesKeys_SInAS_2.5.csv contains the species keys from GBIF.</p><p>DASCO_TaxonHabitats_SInAS_2.5.csv contains habitat information for individual species if available from WoRMS, Fishbase or Sealifebase (used to identify marine species).</p><p>The file DASCO_ListOriginalGBIFData_keys_SInAS_2.5.csv contains the DOIs of the originally downloaded files from GBIF, which provides the basis for the generation of the GBIF part (ie. the DASCO workflow was applied to these data sets from GBIF). Note that OBIS does not provide a DOI for downloads, and thus we cannot provide this.</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Parker Solar Probe Reconnection Exhaust Data Set for Close Encounters 4-11

<p>The data set contained in the ascii text files consist of the current sheet (CS) time intervals and the associated parameters of identified magnetic reconnection exhausts across CSs recorded by the Parker Solar Probe within R&lt;0.26 AU from the Sun for close encounters 4-11. A detailed description is available in a manuscript submitted to The Astrophysical Journal by Eriksson et al. [2024] entitled "Parker Solar Probe Observations of Magnetic Reconnection Exhausts in Quiescent Plasmas Near the Sun".</p><p>The ascii data file zenodo_psp_ce_lmn_cs_output_sorted_tctr.txt contains a center time of the optimized Walen relation (tc, column 1) and columns 2-3 show the CS start (CS1) and end (CS2) times. This is followed by the time duration dtcs in seconds (column 4) and whether a SPAN-I instrument survey (s=sf00) or burst (a=af00) data product exist (column 5). The RTN components of the adjacent &nbsp;time averaged magnetic field before (B1_R, B1_T, B1_N) and after (B2_R, B2_T, B2_N) the CS then follows in columns 6-11 and a corresponding magnetic field rotation angle (Bshear1) at the CS (column 12). The CP dL, dM, dN columns 13-15 list the distances covered along the hybrid-LMN coordinate system based in a Cross-Product normal to the CS in terms of the average ion inertial length diavg (column 16), which is based on a non-corrected proton density (Np1+Np2)/2=Npavg (columns 17-19). The subsequent columns 20-28 are the adjacent VL,VM,VN on the two sides of the CS from the MVAB LMN system as well as their averages (VL1+VL2)=VLavg, (VM1+VM2)/2=VMavg and (VN1+VN2)/2=VNavg. This is followed (columns 29-37) by the corresponding VL,VM,VN and averages in the hybrid-LMN system using the cross-product (CP) normal. The adjacent average total magnetic field Btot1 and Btot2 (columns 38-39) and the LMN components of the magnetic field in the MVAB system [BL1, BL2], [BM1, BM2], [BN1, BN2] in columns 40-45 are followed by the corresponding magnetic fields in the CP hybrid-LMN system (columns 46-51). The ratio of the intermediate to minimum eigenvalue ratio is included (column 52) followed by the unit vectors of the MVAB system L=[L_R, L_T, L_N], M=[M_R, M_T, M_N], and N=[N_R, N_T, N_N] in columns 53-61. The times adjacent to the CS used to obtain this set of MVAB eigenvectors is found in columns 62-63 (MVAB UT1 and MVAB UT2). The magnetic fields (B1 and B2) used to obtain the local hybrid-LMN system relative these MVAB times relatively farther from the CS are associated with a rotation angle (Bshear2) in column 64. The local hybrid-LMN system from the CP normal is then listed in columns 65-73 &nbsp;LCP=[L_R, L_T, L_N], MCP=[M_R, M_T, M_N], NCP=[N_R, N_T, N_N]. The angle (degrees) between the MVAB N and NCP is listed in column 74 (Nmv*Ncp) followed by the last two columns 75-76 that contain the time interval (s) relative CS1 and CS2 to obtain average plasma data (dext) and magentic field data (bext).</p><p>A letter "m" right before column 1 flags the five events in this list of 236 events when the magnetic field did not change sign across the assumed CS. A letter "d" marks the 10 events associated with two opposite (double) exhausts.</p><p>The ascii data file zenodo_psp_ce_lmn_cs_output_sorted_qtn.txt contains the same current sheet (CS) start (CS1) and end (CS2) times (columns 1-2). This file also lists the adjacent average proton (non-corrected) density in columns 3-4 as in the first (tctr) file which is followed by the average proton temperatures (Tp1 and Tp2 in columns 5-6), solar wind speed (Vtot1 and Vtot2 in columns 7-8) and the hybrid system L-components of the velocity (VL1 and VL2 in columns 9-10). Columns 11-16 contain the minimum and maximum values within each CS of the non-correct proton density, proton temperature and VL component. Column 17 lists the daily median density ratio from the electron density (Ne) and the non-corrected proton density (Np) where Ne is obtained through quasi-thermal noise spectroscopy by Kruparova et al. (2023). Column 18 lists the average of this daily ratio. Column 19 lists a value fc. A value larger than 1.0 indicates that a CS event time period is corrected to a local Ne value using Npfc=(Ne/Np)*Np*fc, where Np is a non-corrected proton density, Ne/Np is the median of the daily Ne/Np ratio.</p><p>The ascii data file zenodo_psp_ce_lmn_cs_output_sorted_positions.txt contains the Parker Solar Probe median of its radial position within each CS1-CS2 time period in both solar radius and astronomical unit.</p><p>The PDF figures psp_ceXX_apj_plots_qtn_final.pdf for XX=04,05,06,07,08,09,10 and 11 contain all CS events with a reconnection exhaust for each close encounter 04-11 on the basis of the agreement with a Walen prediction. Each plot shows the pitch-angle distribution (0-180 degrees) of the supra-thermal energy flux, proton temperature (MK), proton density (black) corrected to a daily median ratio Ne/Np as Npcorr=(Ne/Np)*Np, where Ne (red dots) is obtained from a QTN analysis by Kruparova et al. (2023), L-components of the magnetic field and proton velocity with a red trace showing a Walen prediction to this VL across the CS, the M and N components of the magnetic field with BM shifted by its time-period average, and the M and N components of the proton velocity with VM shifted by its time-period average. The LMN unit vectors of the employed hybrid LMN system are shown below each plot for reference.</p><p>Finally, the eight ~11-day overview plots for each close encounter 4-11 marks the center times (tc) of each confirmed exhaust interval in this study of 231 events (red vertical dotted line) with the five marginal events marked as a black vertical dotted line. The panels from the top show the pitch-angle distribution (0-180 degrees) of the supra-thermal energy flux, proton temperature (MK), magnetic field strength, R-components of B and V, N-components of B and V, and the radial postion of Parker Solar Probe in terms of the solar radius. Here, R and N are two of the three RTN system components of B and V.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

A challenging data set for evaluating part-of-speech taggers

<p>This data set contains 2,227 sentences, with a part-of-speech (POS) tag specified for a single word in the sentence. The data file is a tab-separated text file where each row &nbsp;(after the header row) is formatted as follows:</p><p><i>sentence &lt;TAB&gt; POS tag &lt;TAB&gt; (optional) motivation</i></p><p>Note that, in the sentence (= a string of space-separated characters), the POS-tagged word is indicated by the POS tag in brackets, placed just after the word to which it refers. Example:</p><p><i>The road bends [VERB] to the right . VERB</i></p><p>In this example, the optional motivation is not included, as the tagged word can easily be identified as being a verb.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

4D cone beam computed tomography phantom data set

<p>This data set accompanies the following Medical Physics publication: <a href="https://doi.org/10.1002/mp.14441"><i>Madesta, F., Sentker, T., Gauer, T., &amp; Werner, R. (2020). Self‐contained deep learning‐based boosting of 4D cone‐beam CT reconstruction. Medical Physics, 47(11), 5619-5631</i></a><i>.</i></p><p>It comprises 6 time-resolved (4D) cone-beam computed tomography scans with the following scan configurations:</p><ul><li>4D CBCT Scanner: Varian TrueBeam (the detailed scan geometry and further details can be found in Scan.xml included in each scan)</li><li>Phantom: <a href="https://www.cirsinc.com/products/radiation-therapy/dynamic-thorax-motion-phantom/">Dynamic Thorax Phantom: Model 008A</a></li><li>The following motion patterns are included:<ol><li>SI amplitude of insert: ±10mm, pattern: sin, period: 5.0s</li><li>SI amplitude of insert: ±10mm, pattern: cos**4, period: 5.0s</li><li>SI amplitude of insert: ±10mm, pattern: sin, period: 2.5s</li><li>SI amplitude of insert: ±10mm, pattern: cos**4, period: 2.5s</li><li>SI amplitude of insert: ±10mm, pattern: sin, period: 7.5s</li><li>SI amplitude of insert: ±10mm, pattern: cos**4, period: 7.5s</li></ol></li></ul>

opencc-by-nc-sa-4.0Aug 2020View details →
zenodo44/100

CURE Companion Data Set

<p>Companion dataset to superabsorbent polymer manufacture case study described in paper &ldquo;<span>A framework for early-stage sustainability assessment of innovation projects enabled by weighted sum multi-criteria decision analysis in the presence of uncertainty&rdquo;.</span></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Series production data set for 5-axis CNC milling

<p>The data set described encompasses features extracted from the machine control of a five-axis milling machine across thirteen series of productions. Each production series entails a setup changeover to prepare the machine for another product type. Alongside timestamps and twenty features derived from Numerical Control (NC) variables, the data set includes labels denoting various production phases. These labels, up to 23 in total, are structured around a generalized milling process. Comprising thirteen .csv files, each corresponding to a series production, the dataset was gathered within a production company operating in the contract manufacturing sector. These components are tied to actual series orders within ongoing industrial production.</p> <p>The complete description of the data set is published here: <a href="https://www.mdpi.com/2306-5729/9/5/66" target="_blank" rel="noopener">https://doi.org/10.3390/data9050066</a></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

mouse GSE199308 scRNA data set objects

<p>scRNA data from https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE199308 (Huang et al. 2023), see a detailed description of the study here: https://atlas.gs.washington.edu/mmca_v2/public/about.html</p> <p>Data were downloaded from https://atlas.gs.washington.edu/mmca_v2/public/download.html to create an AnnData (h5ad) file with meta data to be able to analyse with e.g. python scanpy package.</p> <p>If you use this data, please cite Huang et al. 2023.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Maximum Independent Set Satellite Scheduling World Cities Data Set

<h1>Satellite Scheduling World Cities Data Set</h1> <p>The Satellite Scheduling World Cities Data Set is the a set of cities treated as point locations used to simulate a set of image collection tasking requests for AIAA paper "A Maximum Independent Set Method for Scheduling Earth-Observing Satellite Constellations".&nbsp;It provides an open reference and benchmark for the satellite task scheduling problem. This could also be considered as&nbsp;a sparse Maximum Independent Set problem for a generic graph. The requests represent point collects, from which we can compute&nbsp;multiple distinct collection opportunities. The tasking problem is then to select a subset of these collects that it is&nbsp;possible for the spacecraft to feasibly collect in a given time period, subject to constraints on the spacecraft's&nbsp;agility and constraints on only collecting a single collect per request (no duplication of effort).<br><br>The data set is hosted on both <a href="https://github.com/duncaneddy/aiaa-mis-satellite-scheduling-dataset">Github</a> and <a href="../">Zenodo</a>. The Github repository contains the original source data, the associated requests generated from the source data, and scripts to reproduce the scenario files. Zenodo (DOI 10.5281/zenodo) hosts copies of the output Metis graph files and collect data files. Due to the large size of produced files these are not included in the Github repository.</p> <h2>Notes</h2> <p><strong>Notes</strong><br><br>Please note that while the source data and generation methods are identical to the satellite&nbsp;task planning paper it was created for. The specific generated problems do not exactly reproduce the&nbsp;scenario in the paper. Since the original reproduction, updates in upstream software dependencies have changed&nbsp;the output of the generation process (specifically, Earth orientaiton parameter handling libraries). This can be&nbsp;determined by considering the cardinality of the generated collect set.&nbsp;However, these differences are generally small and since the constriant rate is similar, the results should be&nbsp;comparable.</p> <table> <tbody> <tr> <td>Spacecraft Count</td> <td>Orignial Publication Collect Count</td> <td>Reproduction Collect Count</td> </tr> <tr> <td>4</td> <td>59356</td> <td>59624</td> </tr> <tr> <td>6</td> <td>90777</td> <td>91204</td> </tr> <tr> <td>12</td> <td>180008</td> <td>180939</td> </tr> <tr> <td>24</td> <td>359170</td> <td>361519</td> </tr> </tbody> </table> <p><br>This repository also adds additional scenarios for 1, 2, and 36 satellites. Note, the&nbsp;provided scenarios represent the largest 10,000 request data set. Should a smaller request set&nbsp;be desired, the requests should be filtered to the top `x` request based on city population and any&nbsp;collects not associated with those requests should be discarded.</p> <p>Note the Zenodo repository excludes the collect and graph files for the 1 and 2 satellite scenarios to avoid the file limits. These can still be reproduced from the Github source code.</p> <h2>Acknolwedgement</h2> <p>If this data set is used in your research, please cite the following paper</p> <p><a href="https://arc.aiaa.org/doi/abs/10.2514/1.A34931">A Maximum Independent Set Method for Scheduling Earth-Observing Satellite Constellations</a></p> <blockquote> <pre><code>@article{eddy2021maximum, title={A Maximum Independent Set Method for Scheduling Earth-Observing Satellite Constellations}, author={Eddy, Duncan and Kochenderfer, Mykel J}, journal={Journal of Spacecraft and Rockets}, volume={58}, number={5}, pages={1416--1429}, year={2021}, publisher={American Institute of Aeronautics and Astronautics} }</code></pre> </blockquote> <h2>Licensing</h2> <p>The source of the world cities data is from the <a href="https://simplemaps.com/data/world-cities">simplemaps.com</a> website,<br>licensed under the <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License </a>with the specific license found at `./data/worldcities_license.txt`.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

JR100 Expedition 379T Site J1002 beryllium isotope, XRF element count and carbon isotope data sets

<h2>JR100 Expedition 379T Site J1002 beryllium isotope, XRF element count and carbon isotope data sets (Finalised 8th of April 2024)</h2> <h3>How to cite these data:</h3> <p>The full data were published in Sproson <em>et al.</em>, 2024.</p> <p>Sproson AD, Yokoyama Y, Miyairi Y, Aze T, Clementi VJ, Riechelson H, Bova SC, Rosenthal Y, Childress LB &amp; Expedition 379T Scientists. Near-synchronous Northern Hemisphere and Patagonian ice sheet variation over the last glacial cycle. <em>Nature Geoscience</em>&nbsp;<a href="https://doi.org/10.1038/s41561-024-01436-y">https://doi.org/10.1038/s41561-024-01436-y</a> (2024).</p> <h3>Files:</h3> <p><strong>Supplementary Table 1: </strong>Multiple linear regression results between 10Be/9Be ratios and sedimentation rate, K/Ca, Fe/Ca, Al/Ti (this study), Green/Blue (Li <em>et al.</em>, 2022), and Global Mean Sea Level (Lambeck <em>et al.</em>, 2014). The multiple linear regression was calculated using the MATLAB(R) function &ldquo;regress&rdquo;.</p> <p><strong>Supplementary Table 2:&nbsp;</strong> Age-depth model and beryllium isotope measurements for Site J1002. The age-depth model was calculated from radiocarbon dates and oxygen isotope stratigraphy (Li <em>et al.</em>, 2022) using the BIGMACS modelling routine (Lee <em>et al.</em>, 2022). Beryllium-9 and beryllium-10 were measured by Adam D. Sproson by HR-ICP-MS and AMS at the Atmosphere and Ocean Research Institute (Sproson <em>et al.</em>, 2021) and University of Tokyo (Matsuzaki et al., 2007), respectively. 10Be/9Be* ratios were corrected for 10Be paleo-production following von Blanckenburg <em>et al.</em> (2015).&nbsp;</p> <p><strong>Supplementary Table 3:</strong> X-ray Fluorescence Ti, K, Fe, Ca, and Al element counts per second for Site J1002 measured at the Lamont-Doherty Earth Observatory by Vincent J. Clementi.</p> <p><strong>Supplementary Table 4: </strong>Carbon isotope measurements for the benthic foraminifera, U. peregrina, measured at Rutgers University by Vincent J. Clementi.</p> <h3>Format:</h3> <p>Depth (m CCSF-A) = core composite depth below seafloor.</p> <p>Calendar age (kyr BP) = age in thousand years before present.</p> <p>[10Be]reac, [9Be]reac = the concentration of 10Be and 9Be in the reactive phase of marine sediments.&nbsp;</p> <p>Sample ID = expedition sample designation specifying hole (e.g., A), core number (e.g., 1), type (i.e., H), section number (e.g., 1), and then section half (i.e., W).</p> <p>&sigma; = standard deviation.</p> <h3>References:</h3> <p>Lambeck K, Rouby H, Purcell A, Sun Y, Sambridge M. Sea level and global ice volumes from the Last Glacial Maximum to the Holocene. <em>Proceedings of the National Academy of Sciences.</em> 2014;111(43):15296-15303.&nbsp;</p> <p>Lee T, Rand D, Lisiecki LE, Gebbie G, Lawrence CE. Bayesian age models and stacks: Combining age inferences from radiocarbon and benthic &delta;18O stratigraphic alignment. <em>EGUsphere.</em> 2022;2022:1-29.</p> <p>Li C, Clementi VJ, Bova SC, <em>et al.</em> The sediment green‐blue color ratio as a proxy for biogenic silica productivity along the Chilean Margin. <em>Geochemistry, Geophysics, Geosystems</em>. 2022:e2022GC010350.&nbsp;</p> <p>Matsuzaki H, Nakano C, Tsuchiya Y, <em>et al.</em> Multi-nuclide AMS performances at MALT. <em>Nuclear Instruments and Methods in Physics Research Section B: Beam Interactions with Materials and Atoms.</em> 2007;259(1):36-40.&nbsp;</p> <p>Sproson AD, Aze T, Behrens B, Yokoyama Y. Initial measurement of beryllium‐9 using high‐resolution inductively coupled plasma mass spectrometry allows for more precise applications of the beryllium isotope system within the Earth Sciences. <em>Rapid Communications in Mass Spectrometry.</em> 2021;35(8):e9059.&nbsp;</p> <p>Von Blanckenburg F, Bouchez J, Ibarra DE, Maher K. Stable runoff and weathering fluxes into the oceans over Quaternary climate cycles. <em>Nature Geoscience. </em>2015;8(7):538-542.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Data set of a thermal response test on a planar trench collector

<p>This data set provides three different temperature measurements (PT100, fiber optic measurement, thermistors), as well as the determination of the volumetric water content and the bulk electrical conductivity of the subsurface. During the published period, a thermal response test was carried out. On April 28th, at 15:25, the fluid circulation began, and heat injection started at 15:38 on April 28th and continued until May 3rd. The test was conducted with a constant volume flow of 1.00 m&sup3;/h and a constant heat injection rate of 0.88 kW.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

NETTAG+ Data set on adsorption of inorganic (Cu and Pb) and organic (PAHs) pollutants in fishing nets

<p>This data set includes the raw data associated with the article in <em>Marine Pollution Bulletin</em><strong> "</strong>Potential of fishing nets for adsorption of inorganic (Cu and Pb) and organic (PAHs) pollutants"</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Hydrographic gridded data set for the South Brazil Bight and Southern Brazilian Shelf

<p>This dataset includes climatological and seasonal maps, spanning data from 1972 to 2024, across 8 different depth levels: 5, 10, 25, 50, 100, 200, 500, and 1000 dBar, with a spatial resolution of 10 km. The maps were generated using the griddata function with triangulation-based natural neighbor interpolation. The variables included in this dataset are conservative temperature (&deg;C), absolute salinity (g kg⁻&sup1;), neutral density (kg m⁻&sup3;), dissolved oxygen (mL L⁻&sup1;), total alkalinity (&micro;mol kg⁻&sup1;), total dissolved inorganic carbon (&micro;mol kg⁻&sup1;), pH (total scale), partial pressure of carbon dioxide (&micro;atm), nitrate (&micro;mol kg⁻&sup1;), and phosphate (&micro;mol kg⁻&sup1;). The name description of each variable is provided in the readme_DatasetATLAS.txt. The dataset can be directly accessed using Ocean Data View (ODV) software.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Data set-PID2020-118478RB-I00

<p>Datos de reconstrucci&oacute;n de tomograf&iacute;as junto con las porosidades detectadas y sus caracter&iacute;sticas.&nbsp;</p>

opencc-zeroNov 2024View details →
zenodo44/100

Data set for letter "Floquet-Driven Crossover from Density-Assisted Tunneling to Enhanced Pair Tunneling"

<p>The files contain the data depicted in the figures of the article "Floquet-Driven Crossover from Density-Assisted Tunneling to Enhanced Pair Tunneling", arXiv 2404.08482.</p> <p>The format of the data and to which figure it corresponds is described in the file "README.txt".</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

A data set from an extensive experimental benchmark study of the Hell Bridge Test Arena subject to imposed damage

<p>A data set from an extensive experimental benchmark study of the Hell Bridge Test Arena (HBTA), a full-scale steel bridge subject to imposed damage, has been established. The data set includes organized dynamic response and load measurement data of the bridge under different structural state conditions, where the structural state conditions range from an undamaged (reference) state to known damage states. Furthermore, the data set includes acceleration and strain data from the response monitoring and acceleration data from the load monitoring, where a modal vibration shaker is used as an excitation source. The data is collected in one h5-file (hierarchical data format version 5) with a sampling rate of 100 Hz. Signal processing and resampling of the data has been performed according to the description provided in the references below. The data set is now published in this open-access data repository and can be accessed and downloaded freely. As such, the data set provides an important benchmark to the scientific community within bridge damage detection and SHM.</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Data set associated to the manuscript entitled Carbon emissions from inland waters may be underestimated: evidence from European river networks fragmented by drying by López-Rojo et. al

<p>CO2 and CH4 emissions and several associated environmental variables &nbsp;were taken in 6 European drying river networks, in 20 river reaches per river network. The field work was carried across 3 sampling campaigns in 2021, coinciding with 3 hydrological seasons (pre-dry, dry and post-rewetting) to encompass most of the hydrological variability. Each time, measures were taken in the habitats available (flowing water, dry riverbeds, isolated pools).</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

A stakeholder-centered determination of High-Value Data sets: the use-case of Latvia

<p>The data in this dataset were collected in the result of the survey of Latvian society (2021) aimed at identifying high-value data set for Latvia, i.e. data sets that, in the view of Latvian society, could create the value for the Latvian economy and society.<br> The survey is created for both individuals and businesses.<br> It being made public both to act as supplementary data for &quot;Towards enrichment of the open government data: a stakeholder-centered determination of High-Value Data sets for Latvia&quot; paper (author: Anastasija Nikiforova, University of Latvia) and in order for other researchers to use these data in their own work.</p> <p>The survey was distributed among Latvian citizens and organisations. The structure of the survey is available in the supplementary file available (see Survey_HighValueDataSets.odt)</p> <p>***Description of the data in this data set: structure of the survey and pre-defined answers (if any)***<br> &nbsp;&nbsp; &nbsp;1. Have you ever used open (government) data? - {(1) yes, once; (2) yes, there has been a little experience; (3) yes, continuously, (4) no, it wasn&rsquo;t needed for me; (5) no, have tried but&nbsp; has failed}<br> &nbsp;&nbsp; &nbsp;2. How would you assess the value of open govenment data that are currently available for your personal use or your business? - 5-point Likert scale, where 1 &ndash; any to 5 &ndash; very high<br> &nbsp;&nbsp; &nbsp;3. If you ever used the open (government) data,&nbsp; what was the purpose of using them? - {(1) Have not had to use; (2) to identify the situation for an object or ab event (e.g. Covid-19 current state); (3) data-driven decision-making; (4) for the enrichment of my data, i.e. by supplementing them; (5) for better understanding of decisions of the government; (6) awareness of governments&rsquo; actions (increasing transparency); (7) forecasting (e.g. trendings etc.); (8) for developing data-driven solutions that use only the open data; (9) for developing data-driven solutions, using open data as a supplement to existing data; (10) for training and education purposes; (11) for entertainment; (12) other (open-ended question)<br> &nbsp;&nbsp; &nbsp;4. What category(ies) of &ldquo;high value datasets&rdquo; is, in you opinion, able to create added value for society or&nbsp; the economy? {(1)Geospatial data;&nbsp; (2) Earth observation and environment; (3) Meteorological; (4) Statistics; (5) Companies and company ownership; (6) Mobility}<br> &nbsp;&nbsp; &nbsp;5. To what extent do you think the current data catalogue of Latvia&rsquo;s Open data portal corresponds to the needs of data users/ consumers? - 10-point Likert scale,&nbsp; where 1 &ndash; no data are useful, but 10 &ndash; fully correspond, i.e. all potentially valuable datasets are available<br> &nbsp;&nbsp; &nbsp;6. Which of the current data categories in Latvia&rsquo;s open data portals, in you opinion, most corresponds to the &ldquo;high value dataset&rdquo;? - {(1)Foreign affairs;&nbsp; (2) business econonmy; (3) energy; (4) citizens and society; (5) education and sport; (6) culture; (7) regions and municipalities; (8) justice, internal affairs and security; (9) transports; (10) public administration; (11) health; (12) environment; (13) agriculture, food and forestry; (14) science and technologies}<br> &nbsp;&nbsp; &nbsp;7. Which of them form your TOP-3? - {(1)Foreign affairs;&nbsp; (2) business econonmy; (3) energy; (4) citizens and society; (5) education and sport; (6) culture; (7) regions and municipalities; (8) justice, internal affairs and security; (9) transports; (10) public administration; (11) health; (12) environment; (13) agriculture, food and forestry; (14) science and technologies}<br> &nbsp;&nbsp; &nbsp;8. How would you assess the value of the following data categories?<br> &nbsp;&nbsp; &nbsp;8.1. sensor data - 5-point Likert scale, where 1 &ndash; not needed to 5 &ndash; highly valuable<br> &nbsp;&nbsp; &nbsp;8.2. real-time data - 5-point Likert scale, where 1 &ndash; not needed to 5 &ndash; highly valuable<br> &nbsp;&nbsp; &nbsp;8.3. geospatial data - 5-point Likert scale, where 1 &ndash; not needed to 5 &ndash; highly valuable<br> &nbsp;&nbsp; &nbsp;9. What would be these datasets? I.e. what (sub)topic could these data be associated with? - open-ended question<br> &nbsp;&nbsp; &nbsp;10. Which of the data sets currently available could be valauble and useful for society and businesses? - open-ended question<br> &nbsp;&nbsp; &nbsp;11. Which of the data sets currently NOT available in Latvia&rsquo;s open data portal could, in your opinion, be valauble and useful for society and businesses? - open-ended question<br> &nbsp;&nbsp; &nbsp;12. How did you define them? - {(1)Subjective opinion; (2) experience with data; (3) filtering out the most popular datasets, i.e. basing the on public opinion; (4) other (open-ended question)}<br> &nbsp;&nbsp; &nbsp;13. How high could be the value of these data sets value for you or your business? - 5-point Likert scale, where 1 &ndash; not valuable, 5 &ndash; highly valuable<br> &nbsp;&nbsp; &nbsp;14. Do you represent any company/ organization (are you working anywhere)? (if &ldquo;yes&rdquo;, please, fill out the survey twice, i.e. as an individual user AND a company&nbsp; representative) - {yes; no; I am an individual data user; other (open-ended)}<br> &nbsp;&nbsp; &nbsp;15. What industry/ sector does your company/ organization belong to? (if you do not work at the moment, please, choose the last option) - {Information and communication services; Financial and ansurance activities; Accommodation and catering services; Education; Real estate operations; Wholesale and retail trade; repair of motor vehicles and motorcycles; transport and storage; construction; water supply; waste water; waste management and recovery; electricity, gas supple, heating and air conditioning; manufacturing industry; mining and quarrying; agriculture, forestry and fisheries professional, scientific and technical services; operation of administrative and service services; public administration and defence; compulsory social insurance; health and social care; art, entertainment and recreation; activities of households as employers;; CSO/NGO; Iam not a representative of any company<br> &nbsp;&nbsp; &nbsp;16. To which category does your company/ organization belong to in terms of its size? - {small; medium; large; self-employeed; I am not a representative of any company}<br> &nbsp;&nbsp; &nbsp;17. What is the age group that&nbsp; you belong to? (if you are an individual user, not a company representative) - {11..15, 16..20, 21..25, 26..30, 31..35, 36..40, 41..45, 46+, &ldquo;do not want to reveal&rdquo;}<br> &nbsp;&nbsp; &nbsp;18. Please, indicate your education or a scientific degree that corresponds most to you? (if you are an individual user, not a company representative) - {master degree; bachelor&rsquo;s degree; Dr. and/ or PhD; student (bachelor level); student (master level); doctoral candidate; pupil; do not want to reveal these data}</p> <p>***Format of the file***<br> .xls, .csv (for the first spreadsheet only), .odt</p> <p>***Licenses or restrictions***<br> CC-BY</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Data set for the integrated Climate, Land, Energy and Water systems modelling exercise RCLEWs in OSeMOSYS

<p>This dataset refers to the modelling exercise (version01_210616RCLEWs).&nbsp;The dataset contains the OSeMOSYS code used to run the modelling exercise, the model input data,&nbsp;the scenarios model data files, and the results. The code for the results visualization is available at&nbsp;https://github.com/KTH-dESA/teaching-CLEWs_visualization.</p> <p>This is an update of version 01_210827&nbsp;available at:&nbsp;https://doi.org/10.5281/zenodo.5293834</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

WaveFake: A data set to facilitate audio DeepFake detection

<p>The main purpose of this data set is to facilitate research into audio DeepFakes. We hope that this work helps in finding new detection methods to prevent such attempts. These generated media files have been increasingly used to commit <a href="https://www.vice.com/en/article/pkyqvb/deepfake-audio-impersonating-ceo-fraud-attempt">impersonation attempts</a>&nbsp;or <a href="https://www.wired.com/story/telegram-still-hasnt-removed-an-ai-bot-thats-abusing-women/">online harassment</a>. You can find the accompanying code repository on&nbsp;<a href="https://github.com/RUB-SysSec/WaveFake">GitHub</a>.</p> <p>The data set consists of &nbsp;104,885 generated audio clips (16-bit PCM wav).&nbsp; We examine multiple networks trained on two reference data sets. First, the <a href="https://keithito.com/LJ-Speech-Dataset/">LJSpeech</a> data set consisting of 13,100 short audio clips (on average 6 seconds each; roughly 24 hours total) read by a female speaker. It features passages from 7 non-fiction books and the audio was recorded on a MacBook Pro microphone. Second, we include samples based on the <a href="https://sites.google.com/site/shinnosuketakamichi/publication/jsut">JSUT</a> data set, specifically, basic5000 corpus. This corpus consists of 5,000 sentences covering all basic kanji of the Japanese language (4.8 seconds on average; roughly 6.7 hours total). The recordings were performed by a female native Japanese speaker in an anechoic room. Finally, we include samples from a full text-to-speech pipeline (16,283 phrases; 3.8s on average; roughly 17.5 hours total). Thus, our data set consists of approximately 175 hours of generated audio files in total. Note that we do not redistribute the reference data.</p> <p>We included a range of architectures in our data set:</p> <ul> <li><a href="https://arxiv.org/abs/1910.06711">MelGAN</a></li> <li><a href="https://arxiv.org/abs/1910.11480">Parallel WaveGAN</a></li> <li><a href="https://arxiv.org/abs/2005.05106">Multi-Band MelGAN</a></li> <li><a href="http://arxiv.org/abs/2005.05106">Full-Band MelGAN</a></li> <li><a href="https://arxiv.org/abs/2010.05646">HiFi-GAN</a></li> <li><a href="https://arxiv.org/abs/1811.00002">WaveGlow</a></li> </ul> <p>Additionally, we examined a bigger version&nbsp;of MelGAN and include samples from a full TTS-pipeline consisting of a conformer and parallel WaveGAN model.</p> <p><strong>Collection Process</strong></p> <p>For WaveGlow, we utilize the <a href="https://github.com/NVIDIA/waveglow">official implementation</a>&nbsp;(commit 8afb643) in conjunction with the official pre-trained network on <a href="https://pytorch.org/hub/nvidia_deeplearningexamples_waveglow/">PyTorch Hub</a>. We use a popular implementation available on <a href="https://github.com/kan-bayashi/ParallelWaveGAN">GitHub</a>&nbsp;(commit 12c677e) for the remaining networks. The repository also offers pre-trained models. We used the pre-trained networks to generate samples that are similar to their respective training distributions, <a href="https://keithito.com/LJ-Speech-Dataset/">LJ Speech</a> and <a href="https://sites.google.com/site/shinnosuketakamichi/publication/jsut">JSUT</a>. When sampling the data set, we first extract Mel spectrograms from the original audio files, using the pre-processing scripts of the corresponding repositories. We then feed these Mel spectrograms to the respective models to obtain the data set.&nbsp;For sampling the full TTS&nbsp;results, we use the <a href="https://github.com/espnet/espnet">ESPnet</a> project.&nbsp;To make sure the generated phrases do not overlap with the training set, we downloaded the <a href="https://commonvoice.mozilla.org/en/datasets">common voices data set</a> and extracted 16.285 phrases from it.</p> <p>This data set is licensed with a CC-BY-SA 4.0 license.</p> <p>This work was supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany&#39;s Excellence Strategy -- EXC-2092 CaSa&nbsp;-- 390781972.</p>

opencc-by-sa-4.0Jun 2021View details →
zenodo44/100

Data set of 'Adiabatic temperature profile in the mantle, revised'

<p>P-V-T data of the four major mantle minerals, olivine, wadselyite, ringwoodite, and bridgmanite</p> <p>The original data are as follows:</p> <p>Olivine: https://doi.org/10.1016/j.pepi.2008.08.002</p> <p>Wadsleyite: https://doi.org/10.1029/2009GL038107</p> <p>Ringwoodite: https://doi.org/10.1029/2004JB003094</p> <p>Bridgmanite; https://doi.org/10.1029/2009GL039318 https://doi.org/10.1029/2011JB008988</p> <p>The temperatures were recalculated using https://doi.org/10.1016/j.pepi.2019.106348</p> <p>The pressures were recalculated using https://doi.org/10.1029/2011JB008988</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record