Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,298
datasets available to search
ShareScore release 0.9.0
Dataset results
1,298 results for “Archiving”
Data archive for Exploiting radar polarimetry for nowcasting thunderstorm hazards using deep learning
<p>This dataset contains the machine learning training data files, pretrained model weights and results for the paper Exploiting radar polarimetry for nowcasting thunderstorm hazards using deep learning, submitted to Natural Hazards and Earth System Sciences, 2023.</p> <p>The radar dataset can be found at the following Zenodo repository: <a href="https://doi.org/10.5281/zenodo.6325370">https://doi.org/10.5281/zenodo.6325370</a></p> <p>For instructions for using the data, please see the GitHub code repository at <a href="http://github.com/meteoswiss/c4dl-polar">https://github.com/meteoswiss/c4dl-polar</a>. Download all the files here and extract the contents to the following subdirectories in the ML code directory:</p> <ul> <li>Training data (patches_quality-index_2020.zip or patches_*_2020.nc) -> data/2020/</li> <li>Results: (results.zip) -> runs/run*/results/</li> <li>Pretrained models (models_run*) -> runs/run*/</li> </ul>
New Data Types in Data Management and Archiving [Webinar recording]
<p>New Data Types in Data Management and Archiving workshop focused on the management, archiving and access to new types of data (NDTs), i.e. administrative, transactional and social media data. The program consisted of four presentations tackling various issues related to handling the NDTs in data repositories and sharing these data in the community of social researchers. Martin Vávra (CSDA) was speaking about current capacities among CESSDA SPs for handling NDTs, Brian Kleiner (FORS) was talking about the coordinated approach to handling NDTs CESSDA SPs. Yevhen Voronin (GESIS) gave a presentation about social media data sharing in social research and Pascal Jurgens (Johannes Gutenberg University Mainz) was speaking about Social Science in the Embattled Digital Age: Adversarial Creation, Use and Sharing of New Data Types. The speakers’ presentations were followed by the panel discussion, where audience members were encouraged to participate and brought in their own experiences of archivists, data managers and researchers. The event was a part of the CESSDA training activities.</p> <p>The video is available on the <a href="https://www.youtube.com/watch?v=j13GsqwDO2Q">CESSDA Training YouTube channel.</a></p>
Journal and Data Archive Collaboration Forum [online event recording]
<p>The availability of research data underlying articles published in journals is becoming a common practice in scientific communication. The European Commission and other funders of scientific research have set high expectations for scientists towards openness and availability of scientific work and results. Scientific publishers, through journals and scholarly publications are the main point of realising open science in practice.<br> <br> This event was part of the continuous Journals Outreach initiative (<a href="https://www.cessda.eu/Training/Journals-outreach">https://www.cessda.eu/Training/Journals-outreach</a>), bringing together CESSDA service providers (SPs) with Social Science & Humanities Journals. <strong>Its target audiences were publishers, editors, researchers, and CESSDA Service providers. </strong>The event was also an opportunity for publishers/journals to highlight new initiatives in research data services linked to scientific publications.<br> <br> The video is available on<a href="https://www.youtube.com/watch?v=zCKoyzLifkg"> the CESSDA Training YouTube channel</a>.</p>
Ports, Past and Present web archive - All Stories
<p>This .wacz file is the web archive for the story collection at https://portspastpresent.eu/, completed on the 13th of July 2023. It captures the navigation options, user experience and story structure of the Omeka story collection at that time for posterity. For more information, see the attached README.</p>
Machine-learning based lightning nowcasting data archive
<p>This data archive contains the derived data supporting the findings of article "Lightning nowcasting with aerosol-informed machine learning and satellite-enriched dataset". The paper is currently in the preprint version: https://doi.org/10.21203/rs.3.rs-2616886/v1</p> <p>The prediction results in this data archive are generated by various models:</p> <p>1. Current model. The model involves data input of aerosol observations together with meteorological variables and auxiliary datasets, as well as data enrichment by Geostationary Lightning Mapper (GLM). In the demo of the dataset, the year of 2020 is trained and predicted on a cross-validation scheme. </p> <p>2. LMA model. The model acts as the baseline model considering only data label obtained from the ground-based Lightning Mapping Array (LMA), which observes accurate lightning occurrence in limited spatial range.</p> <p>3. No-AOD model. The model acts as the baseline model considering no aerosol observation is utilized during the machine learning process. </p> <p>The model results are demonstrated in a continuous value in 0-1. Trade-offs between Probability of Detection (POD) and False Alarm Ratio (FAR) can be optimized by selection of different thresholds. </p> <p>Other datasets:</p> <p>1. Dataset for training. It is for the public use of machine learning training for the current model and no-AOD model (training input features vary).</p> <p>2. PM2.5 dataset. The real-time spatially continuous and hourly-level PM<sub>2.5</sub> dataset is obtained following a published method by Zeng et al.. In this method, the fundamental in-situ measurements are obtained from Air Quality System (AQS) monitoring network operated by United States Environmental Protection Agency.</p> <p>Reference:</p> <p>Siwei Li, Ge Song, Jia Xing et al. Lightning nowcasting with aerosol-informed machine learning and satellite-enriched dataset, 14 March 2023, PREPRINT (Version 1) available at Research Square [https://doi.org/10.21203/rs.3.rs-2616886/v1]</p> <p>Zeng, Z. et al. Estimating hourly surface PM2. 5 concentrations across China from high-density meteorological observations by machine learning. Atmospheric Research 254, 105516 (2021).</p>
InTheMED WP2 Data Archive - Groundwater Quality Literature Review
<p>The data archive InTheMED_WP2_DS_GWQualityLitReview is part of Task 2.2 “Review and collect the available groundwater quantity and quality data sets in the MED region” and contains a literature review of groundwater quality data collected from various sites in Mediterranean countries. The data includes measurements of different water quality parameters, providing valuable insights into the characteristics of groundwater in different regions. The dataset was compiled from a literature review of published research articles.</p>
Variation of and associations with the depth and evenness of sequencing coverage in a sample of archived plastid genomes
<p>Depth and evenness of sequencing coverage are considered potential indicators of genome assembly quality. In plastid genomics, where new data generation has outpaced the development of suitable assembly quality indicators, these coverage metrics could offer insights into the quality of plastomes of different sizes, structures, or taxonomic origins. However, the typical variation of sequencing depth and evenness among archived plastid genomes, their variability between plastome partitions, and any association with methodological factors have yet to be evaluated. This study explores the variation of sequencing depth and evenness across a sample of publicly accessible plastid genomes and their potential associations with plastome structure, assembly accuracy, and the methodological provenance of the genome data using statistical tests. Our results indicate significant differences in sequencing depth across the four structural partitions as well as between the coding and non-coding sections of the genomes, a significant correlation between sequencing evenness and the number of ambiguous nucleotides, and a significant difference in sequencing evenness between several DNA sequencing platforms. These findings highlight that many publicly accessible plastid genomes are based on sequence data with highly variable sequencing depth and evenness and that this variation is influenced, at least partially, by genome structure and methodological factors.</p>
NEON HQ Soil Archive (Megapit) (repackaging of occurrences published by the NEON Biorepository Data Portal)
This collection contains soil samples collected from the megapit at each terrestrial site (NEON sample class: mgp_perarchivesample). During the construction of all 47 terrestrial field sites, "Megapit" soils were collected from multiple horizons at a single soil pit that was up to 2m deep. These samples serve as a reference of soil physical and chemical conditions at the time the NEON site was constructed. The Megapit Archive is curated at the NEON program headquarters in Boulder, CO. Megapit soil samples are available upon request (https://www.neonscience.org/samples/soil-archive).
Ayres 2019: Quantitative Guidelines for Establishing and Operating Soil Archives (repackaging of occurrences published by the NEON Biorepository Data Portal)
Ayres, E. 2019. Quantitative Guidelines for Establishing and Operating Soil Archives. Soil Science Society of America Journal, 83(4): 973-981. https://doi.org/10.2136/sssaj2019.02.0050
SBC LTER Darwin Core Archive: Kelp Forest Reef Fish Abundance
These data describe the abundance of reef fish as part of the Santa Barbara Coastal LTER program (SBC LTER) to track long-term patterns in kelp forest reef species abundance and diversity. The study began in 2000 in the Santa Barbara Channel, California, USA, and the time series is ongoing and updated approximately annually. Abundances of all taxa of resident kelp forest fish encountered along permanent transects are recorded at nine reef sites located along the mainland coast of the Santa Barbara Channel and at two sites on the north side of Santa Cruz Island. These sites reflect several oceanographic regimes in the channel and vary in distance from sources of terrestrial runoff. In these surveys, fish were counted in either a 40x2m benthic quadrat, or in the water parcel 0-2m off the bottom over the same area. This dataset is formatted as a Darwin Core Archive (DwC-A, occurrence core). All taxa are counted (using an open species list), and abundances are zero-filled for each taxon not encountered. This is a derived data product and less-processed data may be available. See http://sbc.lternet.edu for more information and source data, which may include additional measurements, and http://sbc.marinebon.edu for processing notes.
Ant Assemblages in Hemlock Removal Experiment at Harvard Forest since 2003 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/193/5, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-hfr/118/33. The abstract below was extracted from the Level 0 data package and is included for context: Ants comprise a considerable amount of animal biomass in terrestrial ecosystems and play major roles in ecological processes ranging from seed dispersal to soil turnover. Invasion by the hemlock woolly adelgid will transform late-successional hemlock forests into earlier successional mixed hardwood - white pine forests or red-maple wetlands. Understanding how ant assemblages vary in different habitat types allows for predictions of how hemlock decline could alter the composition of ant assemblages, with implications for a wide range of ecosystem processes. As part of the Hemlock Removal Experiment at the Simes Tract, we annually monitor ant species composition and abundance.
North Temperate Lakes LTER: Pelagic Macroinvertebrate Summary 1983 - current (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/284/3, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-ntl/14/32. The abstract below was extracted from the Level 0 data package and is included for context: This is a summary of dataset NTL 13. Derived data include the mean and standard deviation of the number of each species captured as well as the mean and standard deviation of the density of individuals on both an areal and volumetric basis. Five vertical tows are done at the deepest point of each of the seven primary lakes in the Trout Lake area (Allequash, Big Muskellunge, Crystal, Sparkling, and Trout lakes and bog lakes 27-02 [Crystal Bog], and 12-15 [Trout Bog]) using a 1-mm mesh net with a 1-m wide mouth. On Trout Lake four additional sites are sampled, where depths are approximately at 10 m, 15 m, 20 m, and 25 m respectively, with three tows done at each site. Trout Lake was the only lake sampled in 2020. All samples are taken in darkness. Samples are preserved and counted, yielding numbers caught. These night tows target the large invertebrate planktivore component of the pelagic zooplankton community. Sampling Frequency: annually Number of sites: 11
North Temperate Lakes LTER Pelagic Macroinvertebrate Abundance 1983 - current (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/285/3, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-ntl/13/32. The abstract below was extracted from the Level 0 data package and is included for context: Five vertical tows are done at the deepest point of each of the seven primary lakes in the Trout Lake area (Allequash, Big Muskellunge, Crystal, Sparkling, and Trout lakes and bog lakes 27-02 [Crystal Bog], and 12-15 [Trout Bog]) using a 1-mm mesh net with a 1-m wide mouth. On Trout Lake four additional sites are sampled, where depths are approximately at 10 m, 15 m, 20 m, and 25 m respectively, with three tows done at each site. Trout Lake was the only lake sampled in 2020. All samples are taken in darkness. Samples are preserved and counted, yielding numbers caught. These night tows target the large invertebrate planktivore component of the pelagic zooplankton community. Sampling Frequency: annually Number of sites: 11
Gastropod abundance at Hubbard Brook Experimental Forest, Watershed 1 and West of Watershed 6, 1997-2006 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/263/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-hbr/126/4. The abstract below was extracted from the Level 0 data package and is included for context: Snail and slug abundance were measured for a 10 year period between 1997 - 2006 at three elevations on Watershed 1 as well as in a reference area west of Watershed 6. Watershed 1 received calcium additions as wollastonite (CaSiO3) during the study period. This data set includes counts of snails and and slugs for the entire study period. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
Bonanza Creek Experimental Forest Beetles Per Trap Beginning in 1975 - Kruse (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/251/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-bnz/504/7. The abstract below was extracted from the Level 0 data package and is included for context: This is a more detailed datafile then the previous method of reporting found in the datafile: Bonanza Creek Experimental Forest Bark Beetle Per Trap 1Begining in 1975 - Werner. Starting 2010 it contains counts of all woodboring insects and bark beetles caught in the pheromone baited traps, and retains information at the individual sample level.
Eight Mile Lake Research Watershed, Carbon in Permafrost Experimental Heating Research (CiPEHR): Aboveground plant biomass, 2009-2017. (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/275/6, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-bnz/501/17. The abstract below was extracted from the Level 0 data package and is included for context: The Carbon in Permafrost Experimental Heating Research (CiPEHR) project addresses the following questions: 1) Does ecosystem warming cause a net release of C from the ecosystem to the atmosphere?, 2) Does the decomposition of old C that comprises the bulk of the soil C pool influence ecosystem C loss?, and 3) How do winter and summer warming alone, and in combination, affect ecosystem C exchange? We are answering these questions using a combination of field and laboratory experiments to measure ecosystem carbon balance and radiocarbon isotope ratios at a warming experiment located in an upland tundra field site near Healy, Alaska in the foothills of the Alaska Range. This data set includes aboveground plant biomass from winter warming, summer warming, and control treatment plots at CiPEHR.
CBP01 Variable distance line-transect sampling of bird population numbers in different habitats on Konza Prairie (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/339/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-knz/26/11. The abstract below was extracted from the Level 0 data package and is included for context: Records of bird species based on line transect sampling, giving perpendicular distance of sighting from the transect line on 16 separate transects. Bird surveys were conducted 2-4 times per year in January, April, June, and October for a 29-year period from 1981 to 2009. Transects were designed to determine bird communities and population numbers associated with tallgrass prairie habitats with different experimental treatments (fire frequency, grazed by bison vs. ungrazed), riparian habitats on forest edge, and gallery forests dominated by oak woodland.
CFP01 Fish population on selected watersheds at Konza Prairie (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/340/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-knz/87/8. The abstract below was extracted from the Level 0 data package and is included for context: Fishes were collected by habitat (pool or riffle) at 6 sites in the Kings Creek watershed with a single-pass electrofishing survey with one person operating the electrofisher and two people dipnetting. Collections were made seasonally.
PVC02 Plant species composition on selected watersheds at Konza Prairie (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/342/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-knz/69/18. The abstract below was extracted from the Level 0 data package and is included for context: Canopy coverage and frequency were recorded in 20 circular 10 sq m plots. Six treatments were sampled, three ungrazed and three to grazed by native grazers. In each case one of the three watersheds was unburned, another burned annually in April, the third burned every four years in April. In each treatment two soils were sampled: a lower-slope deep fertile nonrocky soil (tully silty clay loam), and a shallow rocky soil (florence cherty silt loam) on level to gently sloping ridges. In 1983 another ungrazed annual burn area '1c' was added 'both tully and florence soils' because original area '1d' appeared aberrant.
Plant aboveground biomass data: BAC: Biodiversity and Climate (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/124/5, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-cdr/386/8. The abstract below was extracted from the Level 0 data package and is included for context: Climate changes forecast for our region by GCM???s and shifts in biodiversity and composition each have the potential to alter ecosystem functioning; their interactive effects are unknown. The "BAC" experiment is designed to determine the direct and interactive effects of plant species numbers, plant community composition, temperature, and precipitation on 11 productivity, C and N dynamics, stability, and plant, microbe, and insect species abundances in CDR grassland ecosystems.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.