Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Dar es Salaam Land Use and Informal Settlement Data Set
The Dar es Salaam Land Use and Informal Settlement Data Set represents urban land use and consolidation of informal settlements for the years 1982, 1992, 1998, and 2002, in Dar es Salaam, Tanzania. The land use categories are informal settlement areas, planned residential areas, ocean and estuaries, vacant and agriculture lands, and other urban features such as industrial or recreation areas. The data are based on the World Geodetic System spheroid of 1984 and use the Universal Transverse Mercator Zone 37 South projection.
India Village-Level Geospatial Socio-Economic Data Set: 1991, 2001
The India Village-Level Geospatial Socio-Economic Data Set: 1991, 2001 is a compilation of the finest level of administrative boundaries in India (village/town-level) and over 200 socio-economic variables collected during the Indian Census in 1991 and 2001. This data set was developed by digitizing village/town level boundaries from the official analog maps published by the Survey of India for 2001. This data set also utilized tabular data for 1991 and 2001 from the Primary Census Abstract (PCA) and Village Directory (VD) data series of the Indian census. The data are in UTM 44N projection and are distributed primarily as shapefiles. Separate files are provided for each of the 28 states (number of states during 1991 and 2001 census) and combined Union Territories for 1991 and 2001.
Global Roads Open Access Data Set, Version 1 (gROADSv1)
The Global Roads Open Access Data Set, Version 1 (gROADSv1) was developed under the auspices of the CODATA Global Roads Data Development Task Group. The data set combines the best available roads data by country into a global roads coverage, using the UN Spatial Data Infrastructure Transport (UNSDI-T) version 2 as a common data model. All country road networks have been joined topologically at the borders, and many countries have been edited for internal topology. Source data for each country are provided in the documentation, and users are encouraged to refer to the readme file for use constraints that apply to a small number of countries. Because the data are compiled from multiple sources, the date range for road network representations ranges from the 1980s to 2010 depending on the country (most countries have no confirmed date), and spatial accuracy varies. The baseline global data set was compiled by the Information Technology Outreach Services (ITOS) of the University of Georgia. Updated data for 27 countries and 6 smaller geographic entities were assembled by Columbia University's Center for International Earth Science Information Network (CIESIN), with a focus largely on developing countries with the poorest data coverage.
Comprehensive Ocean - Atmosphere Data Set (COADS) LMRF Arctic Subset, 1950 - 1995, Version 1
The Comprehensive Ocean - Atmosphere Data Set (COADS) Long Marine Reports Fixed-Length (LMRF) Arctic subset contains marine surface weather reports for regions north of 65 degrees N from ships, drifting ice stations, and buoys. The COADS LMRF Arctic subset contains data collected over the years 1950 to 1995 and includes the following parameters: air and sea temperature, cloudiness, humidity, and winds. The data are in the form of individual marine reports with a given latitude and longitude.
Global Urban Heat Island (UHI) Data Set, 2013
The Urban Heat Island (UHI) effect represents the relatively higher temperatures found in urban areas compared to surrounding rural areas owing to higher proportions of impervious surfaces and the release of waste heat from vehicles and heating and cooling systems. Paved surfaces and built structures tend to absorb shortwave radiation from the sun and release long-wave radiation after a lag of a few hours. The Global Urban Heat Island (UHI) Data Set, 2013, estimates the land surface temperature within urban areas in degrees Celsius (average summer daytime maximum and average summer nighttime minimum) as well as the difference between those temperatures and the temperatures in surrounding rural areas, defined as a 10km buffer around the urban extent. Urban extents are from SEDAC�s Global Rural-Urban Mapping Project, Version 1 (GRUMPv1), and land surface temperatures are from SEDAC�s Global Summer Land Surface Temperature (LST) Grids, 2013, which are derived from the Aqua Level-3 Moderate Resolution Imaging Spectroradiometer (MODIS) Version 5 global daytime and nighttime Land Surface Temperature (LST) 8-day composite data (MYD11A2). For most regions, the UHI data set provides the average daytime maximum (1:30 p.m. overpass) and average nighttime minimum (1:30 a.m. overpass) temperatures in urban and rural areas, and the urban-rural temperature differences, derived from LST data representing a 40-day time-span during July-August (Julian days 185-224) in the northern hemisphere and January-February (Julian days 001-040) in the southern hemisphere. LST grid cells with missing values resulting from high cloud cover in tropical regions were filled with daytime maximum and nighttime minimum LST values from April-May 2013 in the northern hemisphere and December 2013-January 2014 in the southern hemisphere, where available. Some data gaps remain in areas where data were insufficient (e.g., Central Africa).
Anthropogenic Sulfur Dioxide Emissions, 1850-2005: National and Regional Data Set by Source Category, Version 2.86
The Anthropogenic Sulfur Dioxide Emissions, 1850-2005: National and Regional Data Set by Source Category, Version 2.86 provides annual estimates of anthropogenic global and regional Sulfur Dioxide (SO2) emissions spanning the period 1850-2005 using a bottom-up mass balance method, calibrated to country-level inventory data. It includes emissions by country and by source category (coal, petroleum, biomass combustion, smelting, fuel processing, and other processes). This data set is developed at the Pacific Northwest National Laboratory (PNNL) and the maps are produced at the Center for International Earth Science Information Network (CIESIN). The data and maps created using the data set are distributed by the Columbia University Center for International Earth Science Information Network (CIESIN).
Smoke/Sulfates, Clouds and Radiation Experiment in Brazil (SCAR-B) Data Set Version 5.5
SCAR_B_G8_FIRE data are Smoke/Sulfates, Clouds and Radiation Experiment in Brazil, GOES-8 ABBA Diurnal Fire Product (1995 Fire Season) data.Smoke/Sulfates, Clouds and Radiation - Brazil (SCAR-B) data include physical and chemical components of the Earth's surface, the atmosphere and the radiation field collected in Brazil with an emphasis in biomass burning. SCAR-B, the third SCAR experiment, was completed in September 1995, studied the effects of biomass burning on atmospheric processes and aids in the preparation of new techniques for remote sensing of these processes from space.The objectives for the SCAR mission are: to advance our knowledge of how the physical, chemical and radiative processes in our atmosphere are affected by sulfate aerosol and smoke from biomass burning; to improve our expertise at remotely sensing smoke, water vapor, clouds, vegetation and fires; and to assess the effects of deforestation and biomass burning on tropical landscapes.The Cooperative Institute for Meteorological Satellite Studies (CIMSS) at the University of Wisconsin-Madison has produced diurnal GOES-8 derived fire products for the 1995 fire season (June-October 1995) with version 5.5 of the GOES-8 Automated Biomass Burning Algorithm (ABBA). The diurnal fire products were produced for 1145, 1445, 1745, and 2045 UTC coinciding with peak burning hours.The GOES-8 Automated Biomass Burning Algorithm (ABBA) fire products are derived from Geostationary Operational Environmental Satellite (GOES)-8 imager radiances from bands 1 (visible), 2 (3.9 micron), and 4 (11 micron).
CODATA Catalog of Roads Data Sets, Version 1
The CODATA Catalog of Roads Data Sets, Version 1 contains 367 entries describing national-level road network data sets for 147 countries and four entries describing global data sets. It was produced by the Columbia University Center for International Earth Science Information Network (CIESIN) under the oversight of the CODATA Global Roads Data Development Working Group, and as a contribution to the development of the Global Roads Open Access Data Set (gROADS).
Mars surface image (Curiosity rover) labeled data set version 1
This data set consists of 6691 images spanning 24 classes that were collected by the Mars Science Laboratory (MSL, Curosity) rover by three instruments (Mastcam Right eye, Mastcam Left eye, and MAHLI). These images are the "browse" version of each original data product, not full resolution. They are roughly 256x256 pixels each. We divided the MSL images into train, validation, and test data sets according to their sol (Martian day) of acquisition. This strategy was chosen to model how the system will be used operationally with an image archive that grows over time. The images were collected from sols 3 to 1060 (August 2012 to July 2015). The exact train/validation/test splits are given in individual files. Full-size images can be obtained from the PDS at https://pds-imaging.jpl.nasa.gov/search/ .
Food Insecurity Hotspots Data Set
The Food Insecurity Hotspots Data Set consists of grids at 250 meter (~7.2 arc-seconds) resolution that identify the level of intensity and frequency of food insecurity over the 10 years between 2009 and 2019, as well as hotspot areas that have experienced consecutive food insecurity events. The gridded data are based on subnational food security analysis provided by FEWS NET (Famine Early Warning Systems Network) in five (5) regions, including Central America and the Caribbean, Central Asia, East Africa, Southern Africa, and West Africa. Based on the Integrated Food Security Phase Classification (IPC), food insecurity is defined as Minimal, Stressed, Crisis, Emergency, and Famine.
SIGNAL INTEGRATION AND TRANSCRIPTIONAL REGULATION OF THE INFLAMMATORY PROGRAM ASSOCIATED WITH THE GM-/M-CSF SIGNALING AXIS IN HUMAN MONOCYTES [EXPRESSION DATA SET]
GEO Series GSE123270. Homo sapiens. 9 samples. Type: Expression profiling by array.
Early quantification of systemic inflammatory (INF) and cardiovascular disease (CVD) proteins predicts long-term treatment response to Tofacitinib and Etanercept (INF data set)
GEO Series GSE136434. Homo sapiens. 528 samples. Type: Other.
Disruption of the TFAP2A regulatory domain causes Branchio-Oculo-Facial Syndrome (BOFS) and illuminates pathomechanisms for other human neurocristopathies [4C-seq data set 1]
GEO Series GSE108515. Homo sapiens; Gallus gallus. 13 samples. Type: Other.
E-Predict Training Data Set and Examples
GEO Series GSE2228. Viruses. 56 samples. Type: Expression profiling by array.
REST and Neural Gene Network Dysregulation in iPS Cell Models of Alzheimer’s Disease (Affymetrix NPC data set)
GEO Series GSE117586. Homo sapiens. 10 samples. Type: Expression profiling by array.
Risk SNPs mediated promoter-enhancer switching promotes prostate cancer progression through lncRNA PCAT19 (ChIP-seq data set)
GEO Series GSE112120. Homo sapiens. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
MAQC-II Project: Hamner data set
GEO Series GSE24061. Mus musculus. 88 samples. Type: Expression profiling by array.
Differential gene expression pattern of mouse livers transfected with dual oncogenes (day-7 post-injection data set)
GEO Series GSE109544. Mus musculus. 18 samples. Type: Expression profiling by high throughput sequencing.
Expression profilings of fission yeast strains (RNA-seq data set)
GEO Series GSE104546. Schizosaccharomyces pombe. 14 samples. Type: Expression profiling by high throughput sequencing.
Expression data from ALL patients included in the set used to construct a classification signature (COALL cohort)
GEO Series GSE13425. Homo sapiens. 190 samples. Type: Expression profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.