Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,298
datasets available to search
ShareScore release 0.9.0
Dataset results
1,298 results for “Archiving”
SGS-LTER Standard Production Data: 1983-2008 Annual Aboveground Net Primary Production on the Central Plains Experimental Range, Nunn, Colorado, USA 1983-2008, ARS Study Number 6 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/325/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/700/1. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. The objective of the long-term ANPP study is to monitor long-term net above ground primary production of the shortgrass steppe community by species. There are 6 sites: ridgetop (ridge), midslope (mid), swale, ESA (replicate 1 not 2), Section 25 (SEC 25), and owl-creek (OC). Each site is located in a different landscape position or soil type on the shortgrass steppe and may be grazed or not. Ridgetop, midslope and swale are grazed and are sampled along a catena. Section 25 is grazed and is located in an upload grassland. ESA is an ungrazed upland grassland an is the control from the Ecosystem Stress Area experiment. Owl Creek is ungrazed and is located in the lowland along the owl creek drainage. There are 3 transects with 5 plots in each transect. Plots in the grazed
SGS-LTER Long-Term Monitoring Project: Vegetation Cover on Small Mammal Trapping Webs on the Central Plains Experimental Range, Nunn, Colorado, USA 1999 -2006, ARS Study Number 118 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/326/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/140/17. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Additional information and referenced materials can be found: http://hdl.handle.net/10217/83458. The abundance and diversity of small mammals in shortgrass steppe is strongly influenced by the structure and composition of vegetation. Vegetation structure provides cover from predators and harsh abiotic conditions. Plant species composition affects the types of seeds and herbaceous material available to granivores and herbivores, and influences arthropod populations, which are important prey for the omnivorous species that dominate in shortgrass steppe. Both vegetation structure and plant community composition are sensitive to the availability of precipitation as well as the activity of large mammalian herbivores. In 1999, we began measuring vegetation structure and p
SGS-LTER Long-term Monitoring Project: Spotlight Rabbit Count on the Central Plains Experimental Range, Nunn, Colorado, USA 1994-2006, ARS Study Number 98 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/327/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/136/17. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Additional information and referenced materials can be found: http://hdl.handle.net/10217/83448. Rabbits are the most important small-mammal herbivores in shortgrass steppe, and may significant influence the physiognomy and population dynamics of herbaceous plants and woody shrubs. Rabbits also are the most important prey of mammalian carnivores such as coyotes and large raptors such as golden eagles and great horned owls. Two hares (Lepus californicus, L. townsendii) and one cottontail rabbit (Sylvilagus audubonii) occur in shortgrass steppe. In 1994, we initiated long-term studies to track changes in relative abundance of rabbits on the Central Plains Experimental Range (CPER). On four nights each year (one night each season, usually on new moon nights in January,
SGS-LTER Long-Term Monitoring Project: Small Mammals on Trapping Webs on the Central Plains Experimental Range, Nunn, Colorado, USA 1994 -2006, ARS Study Number 118 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/329/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/137/17. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Additional information and referenced materials can be found: http://hdl.handle.net/10217/83452. Small mammals (rabbits, rodents) are integral components of semiarid ecosystems because of their roles as consumers of plants, seeds and arthropods, as soil disturbance agents, and as food for raptors, snakes and mammalian carnivores. Because of their vagility and intermediate trophic position, populations of small mammals may track changes in vegetation and the abiotic environment that may result from shifts in land-use and other anthropogenic disturbances. However, these populations are variable over space and time, and their response to environmental changes may not be immediately apparent given their behavioral flexibility and relatively long life-spans and generatio
SGS-LTER Ecosystem Stress Area - long-term density dataset following nutrient enrichment stress on the Central Plains Experimental Range in Nunn, Colorado, USA 1975-2011, ARS Study Number 3 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/330/3, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/520/8. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Water, nitrogen, and water-plus-nitrogen at levels beyond the range normally experience by shortgrass steppe communities were applied from 1971 through 1975, plant densities were sampled through 1977, and then sampling resumed in 1982, with sampling frequencies changing from annually to every other year. The initial sampling from 1970 to 1974 showed that the water and water plus nitrogen treatments had the strongest effect on plant community structure, both treatments increased biomass, and exotic weed species were noted on the water plus nitrogen treatment. Later sampling from 1982 to 1991 showed a ten-fold increase in exotic weed species on the water plus nitrogen plots as compared to the controls (Milchunas and Lauenroth 1995), a community change that has persiste
SGS-LTER Ecosystem Stress Area - long-term point-frame (percent basal cover) dataset following nutrient enrichment stress on the Central Plains Experimental Range in Nunn, Colorado, USA 1982-2011, ARS Study Number 3 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/331/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/521/7. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Water, nitrogen, and water-plus-nitrogen at levels beyond the range normally experience by shortgrass steppe communities were applied from 1971 through 1975, plant densities were sampled through 1977, and then sampling resumed in 1982, with sampling frequencies changing from annually to every other year. The initial sampling from 1970 to 1974 showed that the water and water plus nitrogen treatments had the strongest effect on plant community structure, both treatments increased biomass, and exotic weed species were noted on the water plus nitrogen treatment. Later sampling from 1982 to 1991 showed a ten-fold increase in exotic weed species on the water plus nitrogen plots as compared to the controls (Milchunas and Lauenroth 1995), a community change that has persiste
Rabbit Population Dynamics in Chihuahuan Desert Grasslands and Shrublands at the Sevilleta National Wildlife Refuge, New Mexico (1992-present) (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/334/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sev/23/121705. The abstract below was extracted from the Level 0 data package and is included for context: This study explores the population dynamics of black-tail jackrabbits (Lepus californicus) and desert cottontail rabbits (Sylvilagus auduboni) in the grasslands and creosote shrublands of McKenzie Flats, Sevilleta National Wildlife Refuge. The study was initiated in January 1992, and continues quarterly each year. Rabbits are sampled via night-time spotlight transect sampling along the roads of McKenzie Flats once during winter, spring, summer, and fall. The route is 21.5 miles long. Measurements of perpendicular distance of each rabbit from the center of the road are used to estimate densities (number of rabbits per square kilometer) via Program DISTANCE. Results from January 1992 to May 2004 indicated that spring was the period of peak density period, with generally steady declines through the rest of the year until the following spring. Evidence of a long-term "cycle" (e.g., the 11-year-cycle reported for rabbits in the Great Basin Desert) does not appear in the Sevilleta rabbit populations.
Bird Abundances at the Hubbard Brook Experimental Forest (1969-present) and on three replicate plots (1986-2000) in the White Mountain National Forest (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/355/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-hbr/81/7. The abstract below was extracted from the Level 0 data package and is included for context: Bird abundances have been determined from timed censuses, territory maps and nest locations at the Hubbard Brook Experimental Forest from 1969 to the present. This data set includes counts of the number of adult birds (males and females) per 10 ha at HBEF (1969 - present) and on three additional plots within the White Mountain National Forest (1986 - 2000). These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
Virgin Islands National Park: Coral Reef: Population Dynamics: Scleractinian corals (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/357/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/291/2. The abstract below was extracted from the Level 0 data package and is included for context: These data are evidence of the the long-term dynamics of shallow coral reefs along the south coast of St. John from as early as 1987. These data describe coral reef community structure as percent cover based on the analysis of color photographs. All of these data originate from color images of photoquadrats recorded annually (usually in the summer) from as early as 1987. The data falls into three groups. The two groups that are contained in this data package are (1) Tektite & Yawzi and (2) Random sites. The juvenile coral density is packaged separately. Tektite – this is at 14 m depth on the eastern side of Great Lameshur Bay and is the original site of the Tektite man-in-the sea project in 1969; this project marked the birth of the Virgin Islands Ecological Research Station (later the Virgin Islands Environmental Resource Station) that hosts the field component of the project. The reef in this location consists of a single buttress that has remained dominated by Montastraea anularis since the start of the research (1987). These surveys consist of 30 photoquadrats (1 x 1 m) distributed along three, 10 m transects. Yawzi – this is at 9 m depth and is on the western side of Great Lameshur Bay and has been recorded photographically since 1987. This reef also started the study period dominated by Montastraea annularis, but has degraded much more rapidly that the Tektite site. These surveys consist of 30 photoquadrats (1 x 1 m) distributed along three, 10 m transects. Random sit
Aphid Collection Tower Site at KBS at the Kellogg Biological Station, Hickory Corners, MI (2005 to 2013) (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/347/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-kbs/49/25. The abstract below was extracted from the Level 0 data package and is included for context: Survey of migration of soybean aphid and other aphids of economic interest in 10 midwestern States. Aphids are collected using a suction trap. original data source http://lter.kbs.msu.edu/datasets/52
SGS-LTER Long-Term Montioring Project: Arthropod Pitfall Trapping on Small Mammal Trapping Webs on the Central Plains Experimental Range, Nunn, Colorado, USA 1998-2006, ARS Study Number 118 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/328/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/134/17. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Additional information and referenced materials can be found: http://hdl.handle.net/10217/83450. With the exception of heteromyids, eg kangaroo rats and pocket mice, most small rodents in shortgrass steppe are omnivorous. Depending on season, arthropods (insects and arachnids) make up 40-85% of the diet of grasshopper mice and thirteen-lined ground squirrels, the most widespread rodents in northern shortgrass steppe. Small mammals are among the most important predators of ground-dwelling macroarthropods and herbivorous insects provide a direct resource link between weather and plant production. Understanding temporal variability in the abundance of arthropods is central to determining the mechanisms that drive small rodent populations. At present, there are no long-
Ecological Survey of Central Arizona: a survey of key ecological indicators in parcels of residential areas in the greater Phoenix metropolitan area, ongoing since 2010 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/115/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-cap/653/2. The abstract below was extracted from the Level 0 data package and is included for context: The Ecological Survey of Central Arizona (ESCA) is an extensive field survey and integrated inventory designed to capture key ecological indicators of the CAP LTER study area consisting of the urbanized, suburbanized, and agricultural areas of metropolitan Phoenix, and the surrounding Sonoran desert. The survey is conducted every five years at approximately 200 sample plots (30m x 30m) that were located randomly using a tessellation-stratified dual-density sampling design. Study plots cover habitats throughout the CAP LTER study area ranging from native Sonoran desert sites to residential yards to an airport tarmac. In 2010, the survey was expanded to include an assessment of residential parcels overlapping the survey plot at sites in residential areas. Many of the same variables that are measured in the 30m x 30m survey plot are measured in the parcel, including an inventory of perennial plants, and the biovolume of trees. In addition, a detailed assessment of characteristics of the parcel is performed. Investigators interested in data from the broader Ecological Survey of Central Arizona that includes all survey plots should should search the data catalog for 'ecological survey of central arizona' or 'survey 200' to locate those and other data related to the CAP LTER's ESCA.
Ecological Survey of Central Arizona: a survey of key ecological indicators in the greater Phoenix metropolitan area and surrounding Sonoran desert, ongoing since 1999 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/247/3, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-cap/652/3. The abstract below was extracted from the Level 0 data package and is included for context: The Ecological Survey of Central Arizona (ESCA) is an extensive field survey and integrated inventory designed to capture key ecological indicators of the CAP LTER study area consisting of the urbanized, suburbanized, and agricultural areas of metropolitan Phoenix, and the surrounding Sonoran desert. The survey, formerly known as the survey 200 and renamed to ESCA in 2015, is conducted every five years at approximately 200 sample plots (30m x 30m) that were located randomly using a tessellation-stratified dual-density sampling design. Study plots cover habitats throughout the CAP LTER study area ranging from native Sonoran desert sites to residential yards to an airport tarmac. Measurements include an inventory of all plants (identified to the lowest possible taxonomic unit, typically species), plant biovolume, soil coring for physicochemical properties, arthropod sweep-net sampling, photo documentation, and a visual survey of site and area characteristics. The objectives of the survey are to (1) characterize patches in terms of key biotic, physical, and chemical variables, and (2) examine relationships among land use, general plant diversity, native plant diversity, plant biovolume, soil nutrient status, and social-economic indices along an indirect urban gradient. A pilot survey was conducted in 1999, and the first full ESCA was conducted in 2000. The maiden survey in 2000 featured a suite of measurements that were not assessed in later surveys, including data fr
Long-term trends in abundance of Lepidoptera larvae at Hubbard Brook Experimental Forest and three additional northern hardwood forest sites, 1986-2018 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/349/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-hbr/82/8. The abstract below was extracted from the Level 0 data package and is included for context: Numbers and lengths of Lepidoptera larvae (caterpillars, all species) were censused on shrub level foliage at biweekly intervals from late May/early June through late July/early August each year. Measurements were conducted on the Main bird plot in the Hubbard Brook Experimental Forest and on three additional plots within the White Mountain National Forest from 1986-1997.. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
Catalog of NCBI sequence read archive (SRA) data for salamanders at the Hubbard Brook Experimental Forest 2012-2021
This project was designed to describe fine-scale population genetic differentiation of the stream salamander Gryinophilus porphyriticus among five study streams in the Hubbard Brook Experimental Forest. The data are paired with intensive capture-recapture data to assess direct fitness effects of individual genetic diversity, including effects of individual multilocus heterozygosity on stage-specific survival probabilities. This dataset publishes a manifest of the genomic sequence reads submitted to the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA). These samples are published at NCBI under the BioProject ID 1090913 (https://www.ncbi.nlm.nih.gov/bioproject/1090913). The tables here include sample metadata and the NCBI URLs to each sample. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
Archived plant, soil, water and microbial samples at the Kellogg Biological Station, Hickory Corners, MI (1988 to 2016)
Dataset AbstractPlants, soil, water and microbial samples are archived. Small quantities are available for research by contacting: Plant and Soil Samples — Stacey Vanderwulp Water Samples — Steve Hamilton Microbial Samples — Tom Schmidt original data source http://lter.kbs.msu.edu/datasets/55
Long-term fish abundance data for Wisconsin Lakes Department of Natural Resources and North Temperate Lakes LTER 1944 - 2012 (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-ntl/346/6, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-ntl/356/3. The abstract below was extracted from the Level 0 data package and is included for context: This dataset describes long-term (1944-2012) variations in the relative abundance of fish populations representing nine species in Wisconsin lakes. Data were collected by Wisconsin Department of Natural Resource fisheries biologists as part of routine lake fisheries assessments. Individual survey methodologies varied over space and time and are described in more detail by Rypel, A. et al., 2016. Seventy-Year Retrospective on Size-Structure Changes in the Recreational Fisheries of Wisconsin. Fisheries, 41, pp.230-243. Available at: http://afs.tandfonline.com/doi/abs/10.1080/03632415.2016.1160894
Treatment of American National Archives Records of World War II Prisoners of War (0326).
<p>NARA PoW Data W.D. A.G.O. FORM NO. 0326.</p> <p>This deposit contains a dataset relating to persons interned between December 7, 1941 and November 19, 1946, which has been enhanced to make it more accessible to scientists. It is based on information from the U.S. National Archives and Records Administration (NARA), which is unrestricted and available at <a href="https://aad.archives.gov/aad/series-description.jsp?s=644&popup=Y">https://aad.archives.gov/aad/series-description.jsp?s=644&popup=Y</a>, and informs this summary. The NARA 'series' is part of <em>Record Group 389: Records of the Office of the Provost Marshal General</em>. It identifies 79 'places of capture' globally "Using copies of reports from the International Committee of the Red Cross ...". The Scope & Content Note states:</p> <p><em>"This series has information about U.S. military officers and soldiers and U.S. and some Allied civilians who were prisoners of war and internees. The record for each prisoner provides serial number, personal name, branch of service or civilian status, grade, date reported, race, state of residence, type of organization, parent unit number and type, place of capture (theater of war), source of report, status, detaining power, and prisoner of war or civilian internee camp site. Records of prisoners of the Japanese who died also document whether the prisoner was on a Japanese ship that sank or if he or she died during transport from the Philippine Islands to Japan. There are no records for some prisoners of war whose names appear in the lists or cables transmitted to the Office of the Provost Marshal General by the International Committee of the Red Cross."</em></p> <p>The U.S. War Department used punched cards to manage this information, although "The punch card records were transferred to NARA with virtually no agency documentation." According to the Custodial History Note:</p> <p><em>"The U.S. Army transferred punch card records of World War II prisoners of war (POWs) to NARA as a unique series in its 1959 transfer of all of the U.S. Army's Departmental Archives. In 1978 the Veterans Administration borrowed most of the punch card records of repatriated U.S. military personnel for a study of Repatriated U.S. Military Prisoners of War, migrated the data on almost all of the borrowed cards to an electronic format and returned the punch cards and two electronic records data files to NARA. In 1995 NARA migrated the data from almost all of the remaining punch card records to an electronic format and has subsequently preserved all of the records in a single data file."</em></p> <p>It is evident that the organization of this data file assumes access to other information, also accessible in CSV files in the series, in order to interpret detailed information, such as branch of service, grade, parent unit number and detaining power. For example, records appear in the following format:</p> <pre><code class="language-bash">O&745255ABDALLAH EDWARD A 2 LT G1AC 200803413223003620O7222094171035 32214872ABDALLAH JOSEPH T CPL 61INF10230241231100157069802075181087 36336867ABDAY JOSEPH C PVT 81INF10170231611100168069516075181004 </code></pre> <p>constituting a serial number, then a name, then a textual code for rank; followed by a string, (starting G1AC on the first line) which encodes the remaining information. For example, the first digit (G) can be looked up in cl_1279.csv to decode ‘2nd lieutenant’, corroborating in this case the appearance of '2 LT. 'AC' indicates 'armofservicecode: AIR CORPS', but less obviously, 'detainingpower: Germany'; 'race: White' and 'theater: European Theater: Germany'. This single line is the entirety of the information provided per person instance by the NARA series. Users of this potentially valuable resource must develop automation in order to be able to search and employ it effectively; no such tools or specification from which software might be developed immediately is provided.</p> <p>Significantly, this task is hampered by evidence of corruption of the some of the information, which may be due solely to the digitization process mentioned above being applied to the paper records, but possibly with subsequent contribution of fixity effects. NARA documentation does not refer to data integrity issues and, especially since the dataset which NARA provides is large, it may only be during development of automation to employ the series that such issues are discovered. Examples of problems include substitution of characters, such as 'O' replacing '0' and vice-versa; '}' replacing '3' and '&' replacing '8', or less obviously 'L' mis-recognized as '-' and 'II' replacing 'H'.</p> <pre><code class="language-bash">12138003 AREY GERALD J S SG 41AC 2002064123S55}340069802055181033</code></pre> <p>Ideally, access to high resolution scans of the paper documents could be used to address these issues, or external documents. However, checking for completeness of each of the components of a person record enables detection of compromised entries and, where character substitution affects decoding of key information, other contextual information is often available to validate decoding such strings with these characters re-substituted. The larger percentage of strings which already decode plausibly without intervention do not contain incidences of such characters (so there is strong evidence that they are invalid in particular positions.</p> <p>Unfortunately, there is a proportion of digital records with more severe corruption which cannot be addressed without access to scanned imagery of the paper records, for example:</p> <pre><code class="language-bash">O&557875ANDREW THOMAS A 2 LT G1AC 2011094115 70140 6881276AFTEWICZ EDWARD L PVT 81INF10150231321100135069508065181004 6 APLIN -OR-& - 3 1INF102 1 1 0 1 1</code></pre> <p>As of the initial date of this deposit is anticipated that such access will be possible to support further work on this series.</p> <p>The dataset in this deposit does not contain records for which decoding is compromised to the extent that information to populate a basic person schema is incomplete. However, although 36,791 of the 143,374 person records in the NARA series were found to be compromised in some way, 19,624 of those have been substantially decoded and/or repaired and further work is being undertaken to both improve decoding of the 126,207 available here and to retrieve others among the 17,167 which are currently inaccessible.</p> <p>This dataset has been enhanced to present the original NARA 'single data file' as a JSON resource which is more accessible for search and analysis, since each record is document-oriented (containing labels and values for each field, together with provenance information) for example:</p> <pre><code class="language-json">{ "$schema": "https://schemata.hasdai.org/historic-persons/historic-person-entry-v0.0.2.json", "location": [ { "association": "military service", "transcription": "European Theater: Germany" }, { "association": "interred", "transcription": "Stalag 2D Stargard Pomerania, Prussia 53-15" } ], "name": { "familyname": "AARON", "givenname": "JACK", "rank": "SGT", "transcription": "AARON JACK" }, "set": { "id": "https://persons.freizo.org/export/pow/1.0.0", "partof": "10.5281/zenodo.3565392", "title": "WDAGO-0326" }, "source": { "type": "data file" } },</code></pre> <p>The schema employed here serves a specific purpose, in addition to on-going work identifying and correcting errors in the NARA data: it supports work to discover other instances of persons appearing in this NARA series 0326, which also appear in external documentation. For example, in a separate collaborative project with Europa Institute at the University of Basel, a benchmark dataset has been produced based on listings of foreign residents in the Asian Directories and Chronicles, which forms a deposit at <a href="https://doi.org/10.5281/zenodo.2580997">10.5281/zenodo.2580997</a> and employs a compatible schema for the purpose of efficient comparison with this and other datasets. Other schemata could be employed for different purposes—leading to alternate datasets, all derived from series 0326. The full extent of information currently decoded from NARA series 0326 is presented at <a href="https://pow.freizo.org/">https://pow.freizo.org/</a> which provides search facilities by person name, plus interactive filters for person rank, service and theater of conflict.</p>
Dataset for paper "Automatically Identifying Archival-worthy, Software-related Slack Conversations"
<p>This dataset consists of 2000 conversations from 5 programming related Q&A channels, hosted on Slack, and accompanies the paper "Automatically Identifying Archival-worthy, Software-related Slack Conversations". In addition to the text of the conversations, each conversation has been annotated as either archival worthy or not. Our definition of archival-worthiness is:</p> <p><em>"If a conversation contains information that could be useful to other users, whether in the Slack channel or elsewhere, then it should be archived. These conversations have no determinate length and no need for objectivity. A conversation should be archived based on the availability and ease of identifying information that could help a person to gain useful software-related knowledge."</em></p> <p><strong>Data Origin: </strong>Numerous public Slack chat channels (<a href="https://slack.com/">https://slack.com/</a>) have recently become available that are focused on specific software engineering-related discussion topics, e.g., Python Development (<a href="https://pyslackers.com/web/slack">https://pyslackers.com/web/slack</a>). The data reflects a portion of the conversations on public channels related to Python, Clojure, Elm and Racket programming.</p> <p><strong>Data Pre-Processing:</strong> To protect privacy, we replace usernames with fake names, and replace absolute times with relative times (in seconds). The conversations are disentangled from the overall chat stream with each unique <em>thread </em>in the dataset specifying a conversation in the channel. Archival-worthy conversations are marked with 1, while non-archival-worthy with 0.</p>
OpenBiodiv Archive
<p>OpenBiodiv is a Knowledge Graph for Literature-Extracted Linked Open Data in Biodiversity Science [1] and is available via http://openbiodiv.net .</p> <p>This data publication contains a single gzipped rdf/nquads archive of all of the OpenBiodiv graph on 2020-04-20 provided by Mariya Dimitrova and uploaded to Zenodo by Jorrit Poelen.</p> <p>Files:</p> <p>openbiodiv.nq.gz - a gzipped rdf/nquads archive of OpenBiodiv</p> <p>openbiodiv.nq.sha256 - sha256 hash of (uncompressed) rdf/nquads archive.</p> <p> </p> <p>[1] Penev, L.; Dimitrova, M.; Senderov, V.; Zhelezov, G.; Georgiev, T.; Stoev, P.; Simov, K. OpenBiodiv: A Knowledge Graph for Literature-Extracted Linked Open Data in Biodiversity Science. <em>Publications</em> <strong>2019</strong>, <em>7</em>, 38. See also <a href="https://doi.org/10.3390/publications7020038">https://doi.org/10.3390/publications7020038 .</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.