Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

45,411

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

45,411 results for “collections”

Learn how ShareScore rates datasets ↗
edi56/100

Cooperative Alaska Forest Inventory (CAFI): V - Photo Collection 2022-2024

The CAFI is a repeated forest measurement project established in forest stands throughout interior and southcentral Alaska. The CAFI was launched in 1994 and measurements were done at a 5-year interval until 2015. The project was on hiatus between 2016 and 2019 but picked back up again in 2020 and will continue at a 10-year interval. Total of 205 permanent plots have been established and each plot has been measured up to 6 times. The CAFI is the most extensive forest monitoring program, both in spatial and temporal scale, in interior and southcentral Alaska today. This is the collection of photos taken at the sites.

openOpenApr 2025View details →
edi56/100

Mollusc population abundance monitoring: Fall 2020 mid-marsh and creekbank infaunal and epifaunal mollusc abundance based on collections from GCE marsh, monitoring sites 1-10

This data set is the Fall 2020 estimate of infaunal and epifaunal mollusc abundance at the GCE-LTER marsh sites used for population monitoring. Species abundance was determined by hand-collecting all the infaunal and epifaunal molluscs from within quadrats of known area in mid-marsh and creekbank zones (n = 4 quadrats per zone) at all sites. The molluscs were returned to the lab, fixed in ethanol, transferred to and preserved in ethanol, counted and measured (size data is reported separately). The counts were converted to number per square meter. Gastropod species are listed first, followed by bivalve species. Size distribution data for these collections may be found in the GCE-LTER data set INV-GCEM-2107a.

openCC (other)Jul 2022View details →
edi56/100

Mollusc population size distribution monitoring: Fall 2020 mid-marsh and creekbank infaunal and epifaunal mollusc size distributions based on collections from GCE marsh monitoring sites 1-10

This data set is the Fall 2020 report of infaunal and epifaunal mollusc species size distributions at the GCE-LTER marsh sites used for population monitoring. Infaunal and epifaunal molluscs were hand-collected from within quadrats of known area from mid-marsh and creekbank zones (n = 4 quadrats per zone) at all sites. The molluscs were returned to the lab, preserved in ethanol, measured and counted (count data is reported separately). Length of each measurable individual was determined using calipers or an ocular micrometer mounted in a stereomicroscope. Species abundance and density data for these collections may be found in the GCE-LTER data set INV-GCEM-2107. Numbers of individuals of each species in the abundance data file may not correspond exactly to the numbers of individuals in the size data file because some individuals may not have been measureable.

openCC (other)Jul 2022View details →
edi56/100

Nitrogen mineralization potential in soils collected from the Jornada Basin LTER-I transect and extracted at field collection time, 1989

This data package contains nitrogen mineralization data from soils collected along the Jornada Basin LTER (LTER-I) transects in southern New Mexico, USA. These transects are located in a livestock exclosure established in 1982 in the Chihuahuan Desert Rangeland Research Center (CDRRC) and run from the middle of the College Playa up to the foot of Mt. Summerford (2.7 km in length). Prior to the exclosure, the study site was moderately to heavily grazed for the past 100 years. The Treatment transect was treated annually with ammonium nitrate fertilizer (NH4NO3 at 10g N/m2/yr) until 1987. Along each transect, 91 stations, each with a plant intercept line, are spaced at 30 meter intervals. For this dataset, 60 soil samples (total) were collected along the control and fertilized treatment transects and mixed with potassium chloride solution (KCl) on Nov 27, 1989, then filter extracted the following day. The dataset contains a soil moisture correction factor, sample weights, total inorganic nitrogen (NO3+NO2-N), and nitrogen in ammonium (NH4-N) for Week F (field) of nitrogen mineralization potentials. The soil mineralization data complements the biomass harvest measurements that occurred in September 1989 (dataset knb-lter-jrn.210015001). This study is complete.

openCC (other)Dec 2021View details →
edi56/100

Nitrogen mineralization potential in soils collected from the Jornada Basin LTER-I transect and extracted at incubation time 0, 1989

This data package contains nitrogen mineralization data from soils collected along the Jornada Basin LTER (LTER-I) transects in southern New Mexico, USA. These transects are located in a livestock exclosure established in 1982 in the Chihuahuan Desert Rangeland Research Center (CDRRC) and run from the middle of the College Playa up to the foot of Mt. Summerford (2.7 km in length). Prior to the exclosure, the study site was moderately to heavily grazed for the past 100 years. The Treatment transect was treated annually with ammonium nitrate fertilizer (NH4NO3 at 10g N/m2/yr) until 1987. Along each transect, 91 stations, each with a plant intercept line, are spaced at 30 meter intervals. For this dataset, 60 soil samples (total) were collected along the control and fertilized treatment transects and mixed with potassium chloride solution (KCl) on Nov 27, 1989, then filter extracted four days later to give a time = 0 incubation value. The dataset contains a soil moisture correction factor, sample weights, total inorganic nitrogen (NO3+NO2-N), and nitrogen in ammonium (NH4-N) for Week 0 of nitrogen mineralization potentials. The soil mineralization data complements the biomass harvest measurements that occurred in September 1989 (dataset knb-lter-jrn.210015001). This study is complete.

openCC (other)Dec 2021View details →
edi56/100

Dissolved inorganic nutrients including 5 macro nutrients: silicate, phosphate, nitrate, nitrite, and ammonium from water column bottle samples collected between October and April at Palmer Station, 1991 - 2025.

The inorganic plant macronutrients dissolved phosphate, silicate, nitrate, nitrite and ammonium are the major sources of nutrition for phytoplankton growth in seawater (with sunlight and inorganic carbon). Macronutrient distributions reflect the large-scale circulation patterns in the oceans and are useful properties to delineate water masses. Dissolved inorganic nutrients samples are typically collected in every Niskin bottle sample collected at and near Palmer Station, Anvers Island, Antarctica on the Western Antarctic Peninsula. Water samples are collected throughout the water column at stations within the Palmer LTER region (primarily B and E, to 50m and 65m respectively). Beginning in the 2020-2021 season, Station B is no longer sampled. In Antarctic waters, dissolved inorganic macronutrients are seldom depleted to limiting concentrations except during heavy prolonged phytoplankton blooms. This is due to the fact that phytoplankton growth is more often limited by light or iron, and to the short growing season. Water samples are analyzed for dissolved nutrients with recognized standard oceanographic protocols for nutrient autoanalyzers (continuous flow analyzers).

openCC (other)Jan 2026View details →
edi56/100

Dissolved inorganic nutrients including 5 macro nutrients: silicate, phosphate, nitrate, nitrite, and ammonium from water column bottle samples collected during annual cruise along western Antarctic Peninsula, 1991 - 2024.

The inorganic plant macronutrients dissolved phosphate, silicate, nitrate, nitrite and ammonium are the major sources of nutrition for phytoplankton growth in seawater (with sunlight and inorganic carbon). Macronutrient distributions reflect the large-scale circulation patterns in the oceans and are useful properties to delineate water masses. Dissolved inorganic nutrients samples are typically collected in every CTD/Rosette cast performed on the annual LTER cruises along the western Antarctic Peninsula. In Antarctic waters, dissolved inorganic macronutrients are seldom depleted to limiting concentrations except during heavy prolonged phytoplankton blooms. This is due to the fact that phytoplankton growth is more often limited by light or iron, and to the short growing season. Water samples pre-filtered through 47mm GF/F filters upon collection and samples frozen until analysis. Water samples are analyzed for dissolved nutrients with recognized standard oceanographic protocols for nutrient autoanalyzers (continuous flow analyzers).

openCC (other)Jan 2026View details →
edi56/100

Euphausia superba length frequency from zooplankton collected with a 2-m, 700-um net towed from surface to 120 m, aboard Palmer LTER annual cruises off the coast of the Western Antarctic Peninsula, 1993 - 2024.

Euphausia superba standard lengths (SL) were measured at grid stations on the annual LTER cruises along the western Antarctic Peninsula (WAP). Annual cruises take place between late December to early February, except for the NBP21-13 cruise, which was November and December. Krill were collected with a 2x2 meter, 700um mesh net fitted with a flow meter and towed obliquely to 120m.

openCC (other)Apr 2025View details →
zenodo52/100

ATTA Biogeochemistry Data Collection

Datasets collected as part of the collaborative projects studying the impact of leaf cutter ants (Atta cephalotes) on biogeochemical cycling in tropical rainforest soils.

opencc-by-4.0Dec 2019View details →
zenodo52/100

Fe chemical speciation collected using trace metal rosette in the Southern Ocean during the austral summer of 2016/2017, on board the Antarctic Circumnavigation Expedition.

<p><strong>Dataset abstract</strong></p> <p>Fe chemical speciation of filtered seawater data are presented in this dataset, resulting from samples collected from a trace metal rosette on board the Antarctic Circumnavigation Expedition (ACE). During the austral summer of 2016/2017, seawater samples were collected from the Atlantic and Indian Ocean sectors of the Southern Ocean and dissolved Fe concentration, iron-binding organic ligands concentration and the conditional stability constant of Fe&rsquo; are presented here.</p> <p><strong>Dataset contents</strong></p> <ul> <li>ace_fe_chemical_speciation.csv, data file, comma-separated values</li> <li>figure1.pdf, metadata, portable document format</li> <li>data_file_header.txt, metadata, text</li> <li>README.txt, metadata, text</li> </ul> <p><strong>Dataset license</strong></p> <p>This Fe chemical speciation dataset from ACE is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>

opencc-by-4.0Jun 2020View details →
zenodo52/100

Hydrolysable carbohydrate data collected from the trace metal rosette in the Southern Ocean during the austral summer of 2016/2017, on board the Antarctic Circumnavigation Expedition.

<p><strong>Dataset abstract</strong></p> <p>Hydrolysable carbohydrate (referred to as TPZT from the analytical methodology used) is part of the labile pool of dissolved organic carbon that is excreted by most (micro)organisms or released by continental margins/sediments. It is a carbon source for heterotrophic bacteria. These carbohydrates could also potentially bind iron and act as an iron binding ligand.</p> <p>This data is used to explore the nature of iron ligands and relate to biological and chemical oceanography.</p> <p><strong>Dataset contents</strong></p> <ul> <li>ace_hydrolysable_carbohydrates_tpzt_data.csv, data file, comma-separated values</li> <li>ace_hydrolysable_carbohydrates_tpzt_data_visual_summary.png, metadata, portable network graphics</li> <li>README.txt, metadata, text format</li> <li>data_file_header.txt, metadata, text format</li> <li>change_log.txt</li> </ul> <p><strong>Dataset license</strong></p> <p>This hydrolysable carbohydrate dataset from ACE is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p> <p><strong>Change log</strong></p> <p>v1.1 - permissions changed to open access (CC BY 4.0 license) and small changes</p> <ul> <li>add license to README.txt</li> <li>format of data_file_header.txt</li> <li>add Frictionless Data schema files</li> </ul> <p>v1.0 - initial release of dataset</p>

opencc-by-4.0Jun 2019View details →
zenodo52/100

Smartphone sensor data (accelerometer, virtual keyboard) collected in-the-wild by Parkinson's Disease patients and Healthy Controls

<p>For detailed description of the dataset see the relevant <a href="https://www.nature.com/articles/s41598-020-78418-8">journal article</a>.</p> <p>Python code for model inference and training is available <a href="https://github.com/alpapado/deep_pd">here</a>.</p> <p>&nbsp;</p> <p><strong>DESCRIPTION</strong></p> <p>The dataset contains accelerometer recodings and keyboard typing data contributed by Parkinson&#39;s Disease patients and Healthy Controls. Accelerometer data consists of acceleration values recorded during phone calls and typing data consist of virtual keyboard press and release timestamps. The dataset is divided into two parts: the first part, called SData, contains data from a small, medically evaluated, set of users, while the second part, called GData, contains recordings from a large body of users with self-reported PD labels.</p> <p>The dataset is organized into 5 pickle files:</p> <p>1. <strong>imu_sdata.pickle</strong>: Contains the tri-axial accelerometer recordings for the SData part of the dataset in the form of a list of python dictionaries, one for each participating subject. Accelerometer data have been pre-processed to a sampling frequency of 100Hz and come segmented into non-overlapping 5 second windows. Hence, a segment&#39;s dimension will be 500 x 3 samples.</p> <p>Sample Python code for accessing the acceleration data of a subject</p> <pre><code class="language-python">sdata = pickle.load(open('imu_sdata.pickle', 'rb')) subject_list = list(sdata.keys()) ## Data for first subject subject_data = sdata[subject_list[0]] # subject_data is a list of length 4 ## The actual data is in the last element of the list acc_segments = subject_data[-1] num_acc_sessions_for_subject = len(acc_segments) acc_segments_for_first_session = acc_segments[0] acc_segments_for_second_session = acc_segments[1] # ..etc In: print(acc_segments_for_first_session.shape) Out: (3, 500, 3) ## The first accelerometer session for this subject consists of 3 five-second segments. In: print(acc_segments_for_second_session.shape) Out: (8, 500, 3) ## The second accelerometer session for this subject consists of 8 five-second segments.</code></pre> <p>2. <strong>imu_gdata.pickle</strong>: Same layout as imu_sdata.pickle but with data ffrom GData subjects.</p> <p>3. <strong>typing_sdata.pickle</strong>: This files contains the typing data originating from the SData part of the dataset. It is a list of dictionaries with one entry per subject. The typing data are given in the form of concatenated hold time (the time elapsed between press and release of the virtual key) and flight time (the time between releasing a key and press the next) histograms, computed over 10ms bins in the range of [0, 1]s for hold time and [0, 4]s for flight time (an additional bin that contains the values in the (1, +oo) and (4, +oo) intervals is also used). So, the total length of the concatenated histogram is 1000/10 + 1 + 4000/10 + 1 = 502.</p> <p>Sample Python code for accessing the typing data of a subject:</p> <pre><code class="language-python">sdata = pickle.load(open('typing_sdata.pickle', 'rb')) subject_list = list(sdata.keys()) ## Data for first subject subject_data = sdata[subject_list[0]] ## The actual data is in the first element of the list typing_histograms = subject_data[0] num_typing_sessions_for_subject = len(typing_histograms) typing_hist_for_first_session = typing_histograms[0] typing_hist_for_second_session = typing_histograms[1] # ..etc In: print(typing_hist_for_first_session.shape) Out: (502, ) ht_hist = typing_hist_for_first_session[:101] # Hold time histogram of the session ft_hist = typing_hist_for_first_session[101:] # Flight time histogram of the session</code></pre> <p>4. <strong>typing_gdata.pickle</strong>: Same layout as typing_sdata.pickle but with data from GData subjects.</p> <p>5. <strong>subject_metadata.pickle</strong>: A list of dictionaries with one entry per subject containing demographic information. The relevant demographic fields have the following interpretation:<br> &nbsp;&#39;age&#39;: Year of birth,<br> &nbsp;&#39;gender_id&#39;: 0 indicates male, 1 indicates female<br> &nbsp;&#39;healthstatus_id&#39;: 0 indicates PD patient, 1 indicates Healthy with PD family history, 2 indicates Healthy without PD family history</p> <p>In the case of SData subjects, there is also symptom UPDRS scores from one or two medical examinations. These are ncoded in the fields med_eval_1 and med_eval_2.</p> <p>&nbsp;</p> <p><strong>ETHICS &amp; FUNDING</strong></p> <p>The study during which the present dataset was collected is a multi-center study approved in each country available (for more info visit: <a href="http://www.i-prognosis.eu/?page_id=3606">http://www.i-prognosis.eu/?page_id=3606</a>).&nbsp;Informed consent, including permission for third-party access to pseudo-anonymised data, was obtained from all subjects prior to their engagement with the study. The work has received funding from the European Union&#39;s Horizon 2020 research and innovation programme under Grant Agreement No 690494 - i-PROGNOSIS: Intelligent Parkinson early detection guiding novel supportive interventions (<a href="http://www.i-prognosis.eu/">i-prognosis.eu</a>).</p> <p>&nbsp;</p> <p><strong>CORRESPONDANCE</strong></p> <p>Any inquiries regarding this dataset should be adressed to:</p> <p>Mr. Alexandros Papadopoulos (Electrical &amp; Computer Engineer, PhD candidate)</p> <p>Multimedia Understanding Groupmug<br> Department of Electrical &amp; Computer Engineering<br> Aristotle University of Thessaloniki<br> University Campus, Building C, 3rd floor<br> Thessaloniki, Greece, GR54124</p> <p>Tel: +30 2310 996359,&nbsp;996365&nbsp;<br> Fax: +30 2310 996398<br> E-mail: alpapado@mug.ee.auth.gr</p> <p>&nbsp;</p> <p><br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo52/100

Atmospheric profiling data collected from radiosondes in the Southern Ocean in the austral summer of 2016/2017 during the Antarctic Circumnavigation Expedition.

<p><strong>Dataset abstract</strong></p> <p>The data set consists of the vertical profiles of the atmospheric variables measured using radiosondes (i-Met) during the Antarctic Circumnavigation Expedition from November 2016 to April 2017. The data include the raw variables measured directly by the radiosondes and derived parameters: altitude (km), air pressure (mb), air temperature (&ordm;C), relative humidity (%), frostpoint (&ordm;C), potential temperature (&ordm;K), water vapour mixing ratio (ppmv), total column water (mm w.e.), wind speed (m/s) and wind direction (deg).</p> <p><strong>Dataset contents</strong></p> <ul> <li>aceNNN_yyyymmdd, directory <ul> <li>aceNNN_yyyymmdd.csv, data file, comma-separated values</li> <li>aceNNN_yyyymmdd.kml, metadata, XML</li> <li>aceNNN_yyyymmdd.raw, data file, raw, ASCII DOS</li> <li>aceNNN_yyyymmdd.raw_config, metadata, XML</li> <li>aceNNN.de1, metadata, ASCII text format</li> <li>aceNNNflt.dat, data file, ASCII text format</li> <li>aceNNNpre.dat, data file, ASCII text format</li> </ul> </li> <li>plots, directory <ul> <li>Sounding_ACENNN.png, metadata, portable network graphics</li> </ul> </li> <li>data_file_header_csv.txt, metadata, text format</li> <li>data_file_header_dat.txt, metadata, text format</li> <li>data_file_header_launches.txt, metadata, text format</li> <li>README.txt, metadata, text format</li> <li>overview_radiosonde_launches.csv, metadata, comma-separated value</li> </ul> <p>where NNN is the launch number yyyy is the year, mm is the month and dd is the day. Dates are in UTC.</p> <p>json files make up a Frictionless Data package.</p> <p><strong>Dataset citation</strong></p> <p>Please cite this dataset as:</p> <p>Gorodetskaya, I.V., Thurnherr, I., Tsukernik, M., Graf, P., Aemisegger, F., Wernli, H. and Ralph, F.M. (2021). Atmospheric profiling data collected from radiosondes in the Southern Ocean in the austral summer of 2016/2017 during the Antarctic Circumnavigation Expedition. (Version 1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.4382460</p>

opencc-by-4.0Jan 2021View details →
zenodo52/100

InnoVine WP3: 105 phenolic compound quantification of 2014 and 2015 mature grape berries from a core-collection of 279 irrigated and non-irrigated Vitis vinifera cultivars

<p>FP7/311775 InnoVine (Innovation in vineyard): Combining innovation in vineyard management and genetic diversity for a sustainable European viticulture</p> <p>WP3: Exploiting the genetic diversity in grapevine</p> <p>105 phenolic or related compounds, from 2014 and 2015 mature grape berries from a core-collection of 279 irrigated and non-irrigated <em>Vitis vinifera</em> cultivars, were quantified by UPLC-TQ-MRM Mass Spectrometry (Lambert M<em> et al., Molecules</em> <strong>2015</strong>, <em>20</em>(5), 7890-7914; doi:10.3390/molecules20057890 &amp; Pinasseau L <em>et al.</em>, <em>Molecules</em> <strong>2016</strong>, <em>21</em>(10), 1409; doi:10.3390/molecules21101409).</p> <p>3 parameters were added:<br> - water/drought status (delta C13)<br> - sugar content (refractive index, brix degree)<br> - weight of 100 grape berries</p> <p>All plant material was collected at the Vassal repository: French National Grapevine Germplasm Collection, INRA Domaine de Vassal, 34340 Marseillan-Plage, France (Centre de Ressources Biologiques de la Vigne (CRB-Vigne) de Vassal-Montpellier).</p>

opencc-by-4.0May 2017View details →
zenodo52/100

Canterbury Museum (CMNZ) collection insect specimen-plant flower interactions

<p>This dataset compromises insect-plant flower interactions recorded from the entomology collections of Canterbury Museum, New Zealand (CMNZ). All invertebrate records were extracted from the Museum Vernon database. This included field collection metadata indicating if an insect and plant flower interaction had occurred. A large proportion of records are specimens collected as part of research by Richard Primack in the 1970s, from insects collected from flowering inflorences. Data was cleaned using OpenRefine v.3.6.2. Insect and plant names were reconciled using the GlobalNames extension in OpenRefine. Data was prepared for submission to the Global Biotic Interaction (GloBI) network.</p> <p>This data makes part of a paper submission to the Journal of Applied Entolomolgy Call for Papers on neglected insects pollinators. If accepted this publication will be linked to this dataset.</p> <p>This version included name updates for insect species, spreadsheet data used to produce summary statistics for the manuscript submission and the README file.</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

A collection of datasets for software vulnerability detection

<p>This is a collection of datasets that are used for AI-based software vulnerability detection. All the datasets are in the .csv format and each row represents a sample. Each dataset includes a set of functions written in C and the target of each function is either 0 (non-vulnerable) or 1 (vulnerable).</p> <ol> <li><strong>data_C_Lin2017_test.csv:</strong> <ul> <li>Reference paper: <a href="https://dl.acm.org/doi/10.1145/3133956.3138840">Vulnerability Discovery with Function Representation Learning from Unlabeled Projects</a>, 2017.</li> <li>Data source on GitHub: <a href="https://github.com/DanielLin1986/function_representation_learning">https://github.com/DanielLin1986/function_representation_learning</a></li> <li>This dataset includes 44 vulnerable and 577 non-vulnerable functions from the LibPNG project.</li> </ul> </li> <li><strong>data_C_LineVul_test.csv:</strong> <ul> <li>Reference paper: <a href="https://ieeexplore.ieee.org/document/9796256">LineVul: A Transformer-based Line-Level Vulnerability Prediction</a>, 2022.</li> <li>Data source on Hugging Face: <a href="https://huggingface.co/datasets/Partha117/LineVul_Test_Dataset">https://huggingface.co/datasets/Partha117/LineVul_Test_Dataset</a></li> <li>This dataset includes 1055 vulnerable and 17809 non-vulnerable functions.</li> </ul> </li> <li><strong>data_C_PrimeVul_test.csv:</strong> <ul> <li>Reference paper: <a href="https://arxiv.org/abs/2403.18624">Vulnerability Detection with Code Language</a><br><a href="https://arxiv.org/abs/2403.18624">Models: How Far Are We?</a> 2024.</li> <li>Data source on GitHub: <a href="https://github.com/DLVulDet/PrimeVul">https://github.com/DLVulDet/PrimeVul</a></li> <li>From the data source, the primevul_test.jsonl was used to created this dataset.</li> <li>This dataset includes&nbsp;695 vulnerable and 25213 non-vulnerable functions.</li> </ul> </li> <li><strong>data_C_Choi2017_test.csv:</strong> <ul> <li>Reference paper: <a href="https://www.ijcai.org/proceedings/2017/0214.pdf">End-to-End Prediction of Buffer Overruns from Raw Source Code</a><br><a href="https://www.ijcai.org/proceedings/2017/0214.pdf">via Neural Memory Networks</a>, 2017.</li> <li>Data source on GitHub: <a href="https://github.com/mjc92/buffer_overrun_memory_networks">https://github.com/mjc92/buffer_overrun_memory_networks</a></li> <li>From GitHub, all the data in trainnig_100.txt, test_1_100.txt, test_2_100.txt,test_3_100.txt,test_4_100.txt, and corresponding _labels.txt files are combined to create this dataset.</li> <li>This dataset includes 7054 vulnerable and 6946 non-vulnerable functions.</li> </ul> </li> <li><strong>data_C_Devign_test.csv:</strong> <ul> <li>Reference paper: <a href="https://proceedings.neurips.cc/paper_files/paper/2019/file/49265d2447bc3bbfe9e76306ce40a31f-Paper.pdf">Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks</a>, 2019</li> <li>Data source on Hugging Face: <a href="https://huggingface.co/datasets/claudios/code_x_glue_devign">https://huggingface.co/datasets/claudios/code_x_glue_devign</a></li> <li>From Hugging Face, all the data in train, validation, and test are combined to create this dataset.</li> <li>This dataset includes&nbsp;12460 vulnerable and 14858 non-vulnerable functions.</li> </ul> </li> <li><strong>data_C_Ours_{train,test}.csv:</strong> <ul> <li>This dataset is manually collected from projects on GitHub that have registered CVEs into NVD from 2002 to 2023. The 6,766 non-vulnerable code functions are extracted from the <a href="https://dl.acm.org/doi/10.1145/3607199.3607242">DiverseVul dataset</a> to increase the code diversity.&nbsp;</li> <li>This training set includes 5413 vulnerable and 5413 non-vulnerable functions.</li> <li>The test set includes 1353 vulnerable and 1353 non-vulnerable functions.</li> </ul> </li> </ol>

openmit-licenseApr 2024View details →
zenodo52/100

Collections (from American Folklife Center)

<p>Dataset originally created 03/01/2019 UPDATE: Packaged on 04/18/2019 UPDATE: Edited README on 04/18/2019</p> <p>I. About this Data Set This data set is a snapshot of work that is ongoing as a collaboration between Kluge Fellow in Digital Studies, Patrick Egan and an intern at the Library of Congress in the American Folklife Center. It contains a combination of metadata from various collections that contain audio recordings of Irish traditional music. The development of this dataset is iterative, and it integrates visualizations that follow the key principles of trust and approachability. The project, entitled, &ldquo;Connections In Sound&rdquo; invites you to use and re-use this data.</p> <p>The text available in the Items dataset is generated from multiple collections of audio material that were discovered at the American Folklife Center. Each instance of a performance was listed and &ldquo;sets&rdquo; or medleys of tunes or songs were split into distinct instances in order to allow machines to read each title separately (whilst still noting that they were part of a group of tunes). The work of the intern was then reviewed before publication, and cross-referenced with the tune index at www.irishtune.info. The Items dataset consists of just over 1000 rows, with new data being added daily in a separate file.</p> <p>The collections dataset contains at least 37 rows of collections that were located by a reference librarian at the American Folklife Center. This search was complemented by searches of the collections by the scholar both on the internet at https://catalog.loc.gov and by using card catalogs.</p> <p>Updates to these datasets will be announced and published as the project progresses.</p> <p>II. What&rsquo;s included? This data set includes:</p> <p>The Items Dataset &ndash; a .CSV containing Media Note, OriginalFormat, On Website, Collection Ref, Missing In Duplication, Collection, Outside Link, Performer, Solo/multiple, Sub-item, type of tune, Tune, Position, Location, State, Date, Notes/Composer, Potential Linked Data, Instrument, Additional Notes, Tune Cleanup. This .CSV is the direct export of the Items Google Spreadsheet</p> <p>III. How Was It Created? These data were created by a Kluge Fellow in Digital Studies and an intern on this program over the course of three months. By listening, transcribing, reviewing, and tagging audio recordings, these scholars improve access and connect sounds in the American Folklife Collections by focusing on Irish traditional music. Once transcribed and tagged, information in these datasets is reviewed before publication.</p> <p>IV. Data Set Field Descriptions</p> <p>IV</p> <p>a) Collections dataset field descriptions</p> <p>ItemId &ndash; this is the identifier for the collection that was found at the AFC<br>Viewed &ndash; if the collection has been viewed, or accessed in any way by the researchers.<br>On LOC &ndash; whether or not there are audio recordings of this collection available on the Library of Congress website.<br>On Other Website &ndash; if any of the recordings in this collection are available elsewhere on the internet<br>Original Format &ndash; the format that was used during the creation of the recordings that were found within each collection<br>Search &ndash; this indicates the type of search that was performed in order that resulted in locating recordings and collections within the AFC<br>Collection &ndash; the official title for the collection as noted on the Library of Congress website<br>State &ndash; The primary state where recordings from the collection were located<br>Other States &ndash; The secondary states where recordings from the collection were located<br>Era / Date &ndash; The decade or year associated with each collection<br>Call Number &ndash; This is the official reference number that is used to locate the collections, both in the urls used on the Library website, and in the reference search for catalog cards (catalog cards can be searched at this address: https://memory.loc.gov/diglib/ihas/html/afccards/afccards-home.html)<br>Finding Aid Online? &ndash; Whether or not a finding aid is available for this collection on the internet</p> <p>b) Items dataset field descriptions</p> <p>id &ndash; the specific identification of the instance of a tune, song or dance within the dataset<br>Media Note &ndash; Any information that is included with the original format, such as identification, name of physical item, additional metadata written on the physical item<br>Original Format &ndash; The physical format that was used when recording each specific performance. Note: this field is used in order to calculate the number of physical items that were created in each collection such as 32 wax cylinders.<br>On Webste? &ndash; Whether or not each instance of a performance is available on the Library of Congress website<br>Collection Ref &ndash; The official reference number of the collection<br>Missing In Duplication &ndash; This column marks if parts of some recordings had been made available on other websites, but not all of the recordings were included in duplication (see recordings from Philadelphia C&eacute;il&iacute; Group on Villanova University website)<br>Collection &ndash; The official title of the collection given by the American Folklife Center<br>Outside Link &ndash; If recordings are available on other websites externally<br>Performer &ndash; The name of the contributor(s)<br>Solo/multiple &ndash; This field is used to calculate the amount of solo performers vs group performers in each collection<br>Sub-item &ndash; In some cases, physical recordings contained extra details, the sub-item column was used to denote these details<br>Type of item &ndash; This column describes each individual item type, as noted by performers and collectors<br>Item &ndash; The item title, as noted by performers and collectors. If an item was not described, it was entered as &ldquo;unidentified&rdquo;<br>Position &ndash; The position on the recording (in some cases during playback, audio cassette player counter markers were used)<br>Location &ndash; Local address of the recording<br>State &ndash; The state where the recording was made<br>Date &ndash; The date that the recording was made<br>Notes/Composer &ndash; The stated composer or source of the item recorded<br>Potential Linked Data &ndash; If items may be linked to other recordings or data, this column was used to provide examples of potential relationships between them<br>Instrument &ndash; The instrument(s) that was used during the performance<br>Additional Notes &ndash; Notes about the process of capturing, transcribing and tagging recordings (for researcher and intern collaboration purposes)<br>Tune Cleanup &ndash; This column was used to tidy each item so that it could be read by machines, but also so that spelling mistakes from the Item column could be corrected, and as an aid to preserving iterations of the editing process</p> <p>V. Rights statement The text in this data set was created by the researcher and intern and can be used in many different ways under creative commons with attribution. All contributions to Connections In Sound are released into the public domain as they are created. Anyone is free to use and re-use this data set in any way they want, provided reference is given to the creators of these datasets.</p> <p>VI. Creator and Contributor Information</p> <p>Creator: Connections In Sound</p> <p>Contributors: Library of Congress Labs</p> <p>VII. Contact Information Please direct all questions and comments to Patrick Egan via www.twitter.com/drpatrickegan or via his website at www.patrickegan.org. You can also get in touch with the Library of Congress Labs team via LC-Labs@loc.gov.</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Pre-processed (in Detectron2 and YOLO format) planetary images and boulder labels collected during the BOULDERING Marie Skłodowska-Curie Global fellowship

<p>This database contains 4976 planetary images of boulder fields located on Earth, Mars and Moon. The data was collected during the BOULDERING Marie Skłodowska-Curie Global fellowship between October 2021 and 2024. The data was already splitted into train, validation and test datasets, but feel free to re-organize the labels at your convenience.&nbsp;</p> <p>For each image, all of the boulder outlines within the image were carefully mapped in QGIS. More information about the labelling procedure can be found in the following manuscript (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a>). This dataset differs from the previous dataset included along with the manuscript&nbsp;<a href="https://zenodo.org/records/8171052">https://zenodo.org/records/8171052</a>, as it contains more mapped images, especially of boulder populations around young impact structures on the Moon (cold spots). In addition, the boulder outlines were also pre-processed so that it can be ingested directly in YOLOv8.</p> <p>A description of what is what is given in the README.txt file (in addition in how to load the custom datasets in Detectron2 and YOLO). Most of the other files are mostly self-explanatory. Please see previous dataset or manuscript for more information. If you want to have more information about specific lunar and martian planetary images, the IDs of the images are still available in the name of the file. Use this ID to find more information (e.g., M121118602_00875_image.png, ID M121118602 ca be used on https://pilot.wr.usgs.gov/). I will also upload the raw data from which this pre-processed dataset was generated (see <a href="https://zenodo.org/records/14250970">https://zenodo.org/records/14250970</a>).</p> <p>Thanks to this database, you can easily train a Detectron2 Mask R-CNN or YOLO instance segmentation models to automatically detect boulders.&nbsp;</p> <p><strong>How to cite:</strong></p> <p>Please refer to the "how to cite" section of the readme file of <a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth.</a></p> <p><strong>Structure:</strong></p> <pre><code>. └── boulder2024/ ├── jupyter-notebooks/ │ └── REGISTERING_BOULDER_DATASET_IN_DETECTRON2.ipynb ├── test/ │ └── images/ │ ├── &lt;image_name&gt;_image.png │ ├── ... │ └── labels/ │ ├── &lt;image_name&gt;_image.txt │ ├── ... ├── train/ │ └── images/ │ ├── &lt;image_name&gt;_image.png │ ├── ... │ └── labels/ │ ├── &lt;image_name&gt;_image.txt │ ├── ... ├── validation/ │ └── images/ │ ├── &lt;image_name&gt;_image.png │ ├── ... │ └── labels/ │ ├── &lt;image_name&gt;_image.txt │ ├── ... ├── detectron2_inst_seg_boulder_dataset.json ├── README.txt ├── yolo_inst_seg_boulder_dataset.yaml</code></pre> <p>&nbsp;</p> <pre><code>detectron2_inst_seg_boulder_dataset.json</code></pre> <p>is a json file containing the masks as expected by Detectron2 (see <a href="https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html">https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html</a> for more information on the format). In order to use this custom dataset, you need to register the dataset before using it in the training. There is an example how to do that in the jupyter-notebooks folder. You need to have detectron2, and all of its depedencies installed. &nbsp;</p> <pre><code>yolo_inst_seg_boulder_dataset.yaml</code></pre> <p>can be used as it is, however you need to update the paths in the .yaml file, to the test, train and validation folders. More information about the YOLO format can be found here (<a href="https://docs.ultralytics.com/datasets/segment/">https://docs.ultralytics.com/datasets/segment/</a>).</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Ten years (2013-2023) of fish assemblage data collected seasonally with underwater visual surveys on paired artificial and natural reefs

<p>The study of assembly patterns and dynamics of organisms has long remained a foundational theme in ecology. Further, the relationship between assemblages and different habitats can provide important insight on ecological processes and guide management and conservation efforts (e.g., restoration, protected areas). We conducted underwater visual surveys of reef fish assemblages at 14 sites in the eastern Gulf of Mexico, including eight that were paired artificial and natural reefs. By using a paired design, we controlled biotic (e.g., larval supply), abiotic (e.g., depth), and socio variables (e.g., fishing access) to isolate the effect of reef type. Trained scientific SCUBA divers with extensive experience with reef fishes from the broader tropical western Atlantic region conducted two to four 10-minute stationary surveys on the paired reefs each season (i.e., calendar quarters) for 10 years from spring 2013 to spring 2023. We also surveyed six additional artificial reefs from winter 2020 to spring 2023 that lacked natural reef pairs. During each survey, the divers identified and estimated the total lengths of all taxa<strong> </strong>observed within an imaginary cylinder around them. The imaginary cylinders had a radius up to 7.5 meters (depending on horizontal visibility) and extended from the seafloor to the highest visible water above the diver. During the period of study, we conducted a total of 1,349 surveys and counted 544,736 fish that represented 171 taxa (most at the species level). Analyses of these data have revealed habitat-specific heterogeneity of the fish assemblages at both taxonomic and functional trait levels, the importance of herbivory in structuring the benthos, and socio-ecological interactions in the system, among other findings. These data may be useful for other researchers interested in patterns and dynamics of populations and communities, functional traits, taxa-habitat relationships, and for parameterizing statistical, joint distribution, metacommunity, and ecosystem models. In addition, because many of the observed taxa<strong> </strong>are of management concern, they may be useful for researchers interested in fisheries science. The data are free to use, are not copyright restricted, and we ask users to cite this data paper.</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

WageIndicator Collective Agreements Database Dataset with Full Texts and Selected Clauses

<p>Since 2012, the&nbsp;<a href="https://wageindicator.org/">WageIndicator Foundation</a>&nbsp;has maintained a&nbsp;<a href="https://wageindicator.org/cbadatabase">Collective Agreements Database</a>, where the texts of 1600 collective agreements (CBAs) from 61 countries and in 27 languages have been uploaded, coded and annotated. This database is a unique example at global level: collective agreements are documents containing conditions of employment that result from negotiations between independent unions and employers, and their content is often surrounded by an atmosphere of secrecy. Under the&nbsp;<a href="https://sshopencloud.eu/">SSHOC project</a>&nbsp;and with the support of the&nbsp;<a href="https://www.clarin.eu/">CLARIN Research Infrastructure</a>, the agreements have been manually and automatically annotated on several levels: for each agreement, the team answers a series of questions and selects the appropriate piece of text (clause) for each.&nbsp;</p> <p>One of the results of the collective agreements&#39; annotation process is the&nbsp;dataset which is available here and includes all the clauses selected for each variable (WageIndicator_CBADatabase_Selected_Clauses). The full collective agreements&#39; texts are stored in another dataset, also available here (WageIndicator_CBADatabase_Full_Texts_211019). A codebook is also included (210125-wageindicator-cba-codebook.pdf).</p>

opencc-by-4.0Dec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record