Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
TwitCID: a Collection of Data Sets for Studies on Information Diffusion on Social Networks
<p>The TwitCID collection consists of five Twitter datasets which were extracted from the 1 percent of tweets from Twitter API. </p> <p>The Firstweek and Secondweek data set were collected during the first week and second week of January 2017 while the Iphone, Gucci and Galaxy data sets were collected from 21 September 2015 to 31 May 2017 using the keywords “iphone”, “gucci” and “galaxys” respectively. </p> <p>We publish these datasets on behalf of our academic institution – IRIT, France and for the sole purpose of non-commercial research under the license CC BY-NC-SA (Attribution-NonCommercial-ShareAlike). In accordance with Twitter's Terms of Service, we only provide identifiers of tweets. In order to collect the actual tweets in JSON, you could use the script Collect_JSONtweets.py attached.</p> <p>If you would like to use this collection, please cite our paper: </p> <p>Hoang, T. B. N., Mothe, J., & Baillon, M. (2019, September). TwitCID: a collection of data sets for studies on information diffusion on social networks. In <em>International Conference of the Cross-Language Evaluation Forum for European Languages</em> (pp. 88-100). Springer, Cham.</p>
"Effects of forestry on summertime low flows and physical fish habitat in snowmelt-dominant headwater catchments of the Pacific Northwest" -- data sets
<p>These files contain the data used in the analysis and production of graphs reported in a manuscript titled "Effects of forestry on summertime low flows and physical fish habitat in snowmelt-dominant headwater catchments of the Pacific Northwest," by Stefan Gronsdahl, R. Dan Moore, Jordan Rosenfeld, Rich McCleary, Rita Winkler. The paper will be published in the journal Hydrological Processes. The file named "readme.txt" explains the contents of the files.</p>
Code and Data from: Segmenting Root Systems in X-Ray Computed Tomography Images Using Level Sets
<p>This record contains code and data for segmentation using a three-dimensional level-set method, written by Amy Tabb in C++. The record also contains two datasets of root systems in media imaged with X-Ray CT, and the results of running the code on those datasets. The code will also perform a pre-processing task in three-dimensional image sets, and a dataset for that purpose is included as well. This work is a companion to the paper : "Segmenting root systems in X-ray computed tomography images using level sets" (WACV 2018) by the authors or this record, and and open-access version of the paper is here -- https://arxiv.org/abs/1809.06398 . The code is also available from Github: https://github.com/amy-tabb/tabb-level-set-segmentation , using a DOI and stable releases https://doi.org/10.5281/zenodo.3344906.</p> <p>Format of the data:</p> <p>Three input datasets are provided; two for the segmentation functionality of the code, and one to test the pre-processing functionality. The two segmentation sets are the same as were used in the paper, and are CassavaDataset, and SoybeanDataset. The pre-processing set is CassavaSlices. The output set for Soybean is SoybeanResultsJul11. The Cassava result set is large, so I broke it into three compressed folders, CassavaResultsJul12_A, _B, _C. _B is the largest, and only contains the results overwritten on the original X-Ray images. Unless your connection to Zenodo is extremely fast, it will be faster to compute the result than to download it.</p> <p> </p> <p> </p><p> </p><p> </p> <p></p> <p></p>
Data set for "Quantification of amorphous siliceous fly ash in hydrating blended cement pastes by X-ray powder diffraction"
<p>The main data is XRD patterns originally collected as xrdml and converted into rd format.</p> <p>The data set for the manuscript:</p> <p>Quantification of amorphous siliceous fly ash in hydrating blended cement pastes by X-ray powder diffraction</p> <p>Xuerun Li<sup>a</sup>, Ruben Snellings<sup>b</sup> and Karen L. Scrivener<sup>a</sup></p> <p><sup>a</sup>Laboratory of Construction Materials, Swiss Federal Institute of Technology in Lausanne (EPFL), Station 12, CH-1015 Lausanne, Switzerland</p> <p><sup>b</sup>Sustainable Materials Management, Flemish Institute of Technological Research (VITO), Boeretang 200, 2400 Mol, Belgium<br> </p>
Images of apples for the use of the Viola-Jones method. Data set no. 2 - grey scale.
<p>The database contains pictures of apples made at different angles, from different sides and containing different varieties. In this way, two bases of apple images were created (each database contains 1,100 images). This set is data set no. 2 - grey scale: processed images in shades of gray. The photos were prepared for the best possible detection process in the Viola-Jones method. These photo bases with apples can be used to teach machines to recognize specific varieties and count apples.</p>
Data set: Industrial IoT-driven remote path planning
<p>This compressed file contains data from three different experiments during the IIoT-REPLAN experimentation phase (Industrial IoT-drive remote path planning). IIoT-REPLAN was funded by an open call from the H2020 Fed4FIRE+ project.</p> <p>Contents:</p> <p>A. astar.csv<br> This file contains the timestamp of each movement of the Robot and the uncertainty (d) at each specific time that Switch 1 was checked. The first two columns refer to seconds while the third one is a scalar value. The total duration of the experiment is 62.33 sec and the setup of this experiment is the real time application of the Astar Algorithm with the localization being based only on the sensor of the Robot</p> <p><br> B. Dijkstra.csv <br> In this experiment the full functionality of the switching system proposed in this work is highlighted. </p> <p>C. Cloud.csv<br> In this experiment the localization algorithm and the path planning algorithm are always executed on the cloud.</p> <p>In both B,C experiments the values of each column are explained inside the Dijkstra.csv </p> <p>Also, two pictures of singlie vision-based self localization are included. </p> <p>A more detailed exposition on all of the above can be found at <br> github link : https://github.com/maravger/alphabot-ppl</p>
Data set for "Distinct contributions of whisker sensory cortex and tongue-jaw motor cortex in a goal-directed sensorimotor transformation"
<p>Data set for: Mayrhofer JM, El-Boustani S, Foustoukos G, Auffret M, Tamura K, Petersen CCH (2019) Distinct contributions of whisker sensory cortex and tongue-jaw motor cortex in a goal-directed sensorimotor transformation. Neuron https://doi.org/10.1016/j.neuron.2019.07.008</p> <p>There are 2 files in this upload:</p> <p>1. The file named "2019_Mayrhofer_Neuron.pdf" is the Open Access pdf file of the manuscript published in Neuron.</p> <p>2. The file named "Mayrhofer_data_code.zip" (~20 GB) is a zipped version of a folder "Mayrhofer_data_code" (~57 GB), which contains the data analysed in the study along with the Matlab code used to generate the published figures. The analysis code is in a subfolder named "MatlabCode", and the specific code for generating each figure panel is in a sub-subfolder named "Figures_tjM1_paper". When running the code, you need to set the Matlab file path to be "Mayrhofer_data_code". In addition, you should add the folder "Mayrhofer_data_code" with subfolders in Matlab "Set Path". The figures will be saved in a subfolder named "Figures". Some parts of the code rely upon previous results, and need to be executed sequentially in the order of the figure panels in the journal publication.</p>
Data Sets and Prediction Models Created using MLP, RNN and LSTM
<p>Binary Data Sets</p> <p>- 2018tbi219_shuffled3.csv and 2018tbi219_shuffled5.csv</p> <p>- these are stratified data sets that were able to produce prediction models with high prediction rates.</p> <p>Models.zip</p> <p>- this zip file contains the prediction models created using Keras Deep Learning Algorithms: MLP, RNN and LSTM</p> <p>Model Creation - Python Code Snippets</p> <p>- Code snippets for creating the prediction models</p>
Supplementary material to the manuscript: Regionalised Heat Demand and Power-To-Heat Capacities in Germany - An Open Data Set for Assessing Renewable Energy Integration
<p>This is the supplementary material for the manuscript:</p> <p>"Regionalised Heat Demand and Power-To-Heat Capacities in Germany - an Open Data Set for Assessing Renewable Energy Integration"</p> <p>Article DOI: <a href="https://doi.org/10.1016/j.apenergy.2019.114161">https://doi.org/10.1016/j.apenergy.2019.114161</a></p> <p>Open access preprint: <a href="https://arxiv.org/abs/1912.03763">https://arxiv.org/abs/1912.03763</a></p> <p> </p> <p><strong>DESCRIPTION OF THE DATASET AND LICENSES:</strong></p> <p>The subdirectory "04_results" contains the regionalised heat demand an power-to-heat capacity data on administrative district level (NUTS-3) for Germany. The subdirectories "01_census_special_evaluation_data" and "02_other_input_data" contain the utilised input data. The subdirectory "03_code" contains the developed and applied source code.</p> <p>The data in this repository are provided under open source licenses. For license information and other general information on the supplementary material, refer to the LICENSE files and README files in the respective subdirectories.</p> <p>For a detailed description of the approach developed by the author, the input data used and the generated results, refer to the manuscript "Regionalised Heat Demand and Power-To-Heat Capacities in Germany - an Open Data Set for Assessing Renewable Energy Integration".</p> <p><strong>METADATA:</strong></p> <p>Sector: Residential Buildings – Space Heating and Domestic Hot Water</p> <p>Geographical scope: Germany</p> <p>Geographical resolution: Administrative districts (NUTS-3)</p> <p>Temporal scope: 2011, three scenarios for 2030</p> <p>Temporal resolution: 15min</p> <p> </p> <p><strong>UNITS:</strong></p> <p>In the final results folders (04_results/01_installed_heating_p2h_capacity; 04_results/02_daily_time_series; 04_results/03_yearly_time_series) the units of the data are indicated in the file names or the column names, e.g. by "in_MW". In case of unit indication in the file name, the unit refers to all columns in the file.</p> <p>In the intermediate results folder (04_results/00_sql_tables_exported_to_csv) all units referring to power are "kW" and all units referring to energy are "kWh".</p> <p><strong>NEWS AND CONTACT:</strong></p> <p>This dataset will be used as part of the <a href="https://wiki.openmod-initiative.org/wiki/Region4FLEX">region4FLEX model</a>. We are currently enhancing the data by temporally and spatially resolved COP time series and determining load shifting potentials. If you wish to receive news or have general questions please contact: wilko.heitkoetter@dlr.de. </p>
King Louie: DBMS Availability Evaluation Data Sets
<p>These data sets provide all availability measurements as accompanying material for the research paper <strong><em>King</em> </strong><em><strong>Louie: Reproducible Availability Benchmarking of Cloud-hosted DBMS</strong> </em>that is presented in the 35th ACM/SIGAPP Symposium on Applied Computing (SAC ’20), March 30-April 3, 2020, Brno, Czech Republic.</p> <p> </p> <p> </p>
Shear Wave Splitting and Mantle Flow beneath Alaska Data Set
<p>Entire data set for the (under review) publication "Shear Wave Splitting in Alaska."</p> <p>McPherson_S1_Station_Info is a table that contains the following columns (with header row): Station Name, Network, Latitude (Deg), Longitude (Deg). This is a table of all the seismic stations in Alaska and western Canada that we downloaded data from. Only stations that were active from Jan 1, 2010, to Aug 18, 2017 are included.</p> <p>McPherson_S2_Event_Info is a table that contains the following columns (with header row): Julian Date, Origin Time, Latitude (Deg), Longitude (Deg), Depth (km), Magnitude (Mw). This is a table of all the seismic events that occurred between Jan 1, 2010, to Aug 18, 2017 within the distance range 80 to 140 degrees from a station, over moment magnitude 5.</p> <p>McPherson_S3_Results_Info is a table that contains the following columns (with header row): Station Name, Back Azimuth (Deg), Distance (Deg), Fast Direction (Deg), Lower Bound (Deg), Upper Bound (Deg), Time Difference (sec), Lower Bound (sec), Upper Bound (sec), Julian Date, Origin Time. This table contains all of the minimum energy method (Silver & Chan, 1991) results that are displayed in Figures 4, 6-12 of the paper under review.</p> <p>McPherson_S4_Nulls_Info is a table that contains the following columns (with header row): Station Name, Back Azimuth (Deg), Distance (Deg), Julian Date, Origin Time. This tables contains all the null results displayed in Figure 5 of the paper under review.</p>
Mowgli: DBMS Performance & Scalability Evaluation Data Sets
<p>These data sets contain the performance and scalability evaluation data created by the <a href="https://omi-gitlab.e-technik.uni-ulm.de/mowgli/getting-started">Mowgli</a> framework for Apache Cassandra and Couchbase, operated on a private Openstack and the Amazon EC2 cloud.</p>
Aalto-1/RADMON data set 2017/2018
<p><strong>General information:</strong></p> <ul> <li>For a description of the Aalto-1 mission, see Kestilä et al. (2013) and Praks et al. (2018)</li> <li>For the RADMON instrument and its calibration, see Peltonen et al. (2014) and Oleynik et al. (2020)</li> <li>For this dataset description, see Gieseler et al. (2020)</li> <li>If you use this data set, please give a proper citation to Gieseler et al. (2020). e.g.: <ul> <li><em>Gieseler, J., Oleynik, P., Hietala, H., Vainio, R., Hedman, H.P., Peltonen, J., Punkkinen, A., Punkkinen, R., Säntti, T., Hæggström, E., Praks, J., Niemelä, P., Riwanto, B., Jovanovic, N., Mughal, M.R., 2020. Radiation Monitor RADMON aboard Aalto-1 CubeSat: First results. Advances in Space Research, 66 (1), 52–65</em>, <em><a href="https://doi.org/10.1016/j.asr.2019.11.023">doi:10.1016/j.asr.2019.11.023</a>, <a href="https://arxiv.org/abs/1911.07586">arXiv:1911.07586</a></em></li> </ul> </li> <li><strong>For omnidirectional averaged fluxes use combination of multiple measurements intervals (e.g., for the same coordinates) </strong><strong>with total integration times of at least 30 minutes</strong><strong> (for more details, see Gieseler et al., 2020)</strong></li> <li>Data gaps in the fluxes are indicated as empty entries, while zeros denote zero counts in the measurement interval</li> <li>Empty entries in AACGM coordinates and L parameter indicate not-defined values</li> </ul> <p> </p> <p><strong>Data file columns description:</strong></p> <ul> <li>time: end time of measurement interval in YYYY-MM-DD HH:MM:SS</li> <li>e2 - e4: integral electron fluxes in (cm<sup>2</sup> s sr)<sup>-1</sup> (see below for energies)</li> <li>i1 - i4: integral proton fluxes (cm<sup>2</sup> s sr)<sup>-1</sup> (see below for energies)</li> <li>i5: differential proton flux (cm<sup>2</sup> s sr MeV)<sup>-1</sup> (see below for energies)</li> <li>interval: integration time of measurement interval (usually 15s)</li> <li>lat: geographic latitude of Aalto-1 </li> <li>lon: geographic longitude of Aalto-1 </li> <li>altitude: altitude above sea level (R<sub>E</sub> = 6371.0 km) of Aalto-1 (in m)</li> <li>mag_lat: geomagnetic latitude of Aalto-1</li> <li>mag_lon: geomagnetic longitude of Aalto-1</li> <li>aacgmv2_lat: altitude-adjusted corrected geomagnetic latitude of Aalto-1</li> <li>aacgmv2_lon: altitude-adjusted corrected geomagnetic longitude of Aalto-1</li> <li>aacgmv2_mlt: magnetic local time based on the AACGM coordinates</li> <li>L: McIlwain L parameter</li> </ul> <p>See Gieseler et al. (2020) for descriptions of the different variables and its derivations.</p> <p> </p> <p><strong>Channel details:</strong></p> <p>Electrons:</p> <ul> <li>e2<br> cutoff energy: 1.51 ± 0.1 MeV<br> geometric factor: 0.0108 ± 0.0005 cm<sup>2</sup> s sr</li> <li>e3<br> cutoff energy: 3.1 ± 0.2 MeV<br> geometric factor: 0.0160 ± 0.0005 cm<sup>2</sup> s sr</li> <li>e4<br> cutoff energy: 6.0 ± 0.7 MeV<br> geometric factor: 0.0119 ± 0.0008 cm<sup>2</sup> s sr</li> </ul> <p>Protons:</p> <ul> <li>i1<br> cutoff energy: 10.4 ± 0.3 MeV<br> geometric factor: 0.0228 ± 0.0004 cm<sup>2</sup> s sr</li> <li>i2<br> cutoff energy: 18.5 ± 0.7 MeV<br> geometric factor: 0.0256 ± 0.0009 cm<sup>2</sup> s sr</li> <li>i3<br> cutoff energy: 23.7 ± 1.8 MeV<br> geometric factor: 0.0219 ± 0.0011 cm<sup>2</sup> s sr</li> <li>i4<br> cutoff energy: 29 ± 4 MeV<br> geometric factor: 0.0187 ± 0.0014 cm<sup>2</sup> s sr</li> <li>i5<br> effective energy: 42 ± 5 MeV<br> differential geometric factor: 0.783 ± 0.09 cm² cm<sup>2</sup> s sr MeV</li> </ul> <p>Note: The channel i5 is the only one yielding differential flux (instead of integral flux), with an energy range of 40-80 MeV, and an effective energy of 42 MeV. Its differential geometric factor is given in units of (cm<sup>2</sup> sr MeV)<sup>-1</sup>.</p> <p>General:</p> <ul> <li>See OIeynik et al. (2020) for more details, e.g., on the derivation of bowtie cutoff energies and possible proton contamination in electron channels</li> <li>Energies are bowtie cutoff energies for integral channels except for proton channel i5 for which the effective energy is given</li> <li>Uncertainties for energies and geometric factors are 95% (two sigmas)</li> <li>All channels yield integral fluxes in units of (cm<sup>2</sup> s sr)<sup>-1</sup> except for proton channel i5 which gives differential flux in (cm<sup>2</sup> s sr MeV)<sup>-1</sup></li> </ul>
Test data sets: File A
I am getting an "Entity too large" error when trying to use this file as File A, with File A taxonID: EOL-000002128790. The Encyclopedia of Life (EOL, eol.org) aggregates biodiversity information from more than 400 sources and provides access to the data through taxon pages, visual query and application programming interfaces. Scientific names are essential elements of the data integration infrastructure, but their shortcomings as key identifiers are well documented (Patterson et al., 2016). Complex automated workflows and continuous manual curation are required to address idiosyncrasies of source taxonomies, variation in data quality, and conflicting taxonomic opinions. To achieve a harmonized taxonomic view of EOL content, names from data sources are mapped to a dynamic reference hierarchy ([see current version here](<p></p>https://opendata.eol.org/dataset/tram-807-808-809-810-dh-v1-1/resource/00adb47b-57ed-4f6b-8f66-83bfdb5120e8)) using an algorithm that leverages canonical name strings, hierarchical information (ancestry, descendants), taxonomic ranks, synonym data, and author strings. Names that cannot be associated with a reference taxon are still accessible, but their unmapped status excludes them and any associated content from certain core EOL functions. For more information about the EOL taxonomy, see [EOL Dynamic Hierarchy](<p></p>https://eol.org/docs/eol-dynamic-hierarchy)
Test data sets: File B
I am getting an "Entity too large" error when trying to use this file as File B, with File B taxonID: BJ5N3. The Encyclopedia of Life (EOL, eol.org) aggregates biodiversity information from more than 400 sources and provides access to the data through taxon pages, visual query and application programming interfaces. Scientific names are essential elements of the data integration infrastructure, but their shortcomings as key identifiers are well documented (Patterson et al., 2016). Complex automated workflows and continuous manual curation are required to address idiosyncrasies of source taxonomies, variation in data quality, and conflicting taxonomic opinions. To achieve a harmonized taxonomic view of EOL content, names from data sources are mapped to a dynamic reference hierarchy ([see current version here](<p></p>https://opendata.eol.org/dataset/tram-807-808-809-810-dh-v1-1/resource/00adb47b-57ed-4f6b-8f66-83bfdb5120e8)) using an algorithm that leverages canonical name strings, hierarchical information (ancestry, descendants), taxonomic ranks, synonym data, and author strings. Names that cannot be associated with a reference taxon are still accessible, but their unmapped status excludes them and any associated content from certain core EOL functions. For more information about the EOL taxonomy, see [EOL Dynamic Hierarchy](<p></p>https://eol.org/docs/eol-dynamic-hierarchy)
Test data sets: Test data set for "Request Entity Too Large" exception
The Encyclopedia of Life (EOL, eol.org) aggregates biodiversity information from more than 400 sources and provides access to the data through taxon pages, visual query and application programming interfaces. Scientific names are essential elements of the data integration infrastructure, but their shortcomings as key identifiers are well documented (Patterson et al., 2016). Complex automated workflows and continuous manual curation are required to address idiosyncrasies of source taxonomies, variation in data quality, and conflicting taxonomic opinions. To achieve a harmonized taxonomic view of EOL content, names from data sources are mapped to a dynamic reference hierarchy ([see current version here](<p></p>https://opendata.eol.org/dataset/tram-807-808-809-810-dh-v1-1/resource/00adb47b-57ed-4f6b-8f66-83bfdb5120e8)) using an algorithm that leverages canonical name strings, hierarchical information (ancestry, descendants), taxonomic ranks, synonym data, and author strings. Names that cannot be associated with a reference taxon are still accessible, but their unmapped status excludes them and any associated content from certain core EOL functions. For more information about the EOL taxonomy, see [EOL Dynamic Hierarchy](<p></p>https://eol.org/docs/eol-dynamic-hierarchy)
Test data sets: COLTest
The Encyclopedia of Life (EOL, eol.org) aggregates biodiversity information from more than 400 sources and provides access to the data through taxon pages, visual query and application programming interfaces. Scientific names are essential elements of the data integration infrastructure, but their shortcomings as key identifiers are well documented (Patterson et al., 2016). Complex automated workflows and continuous manual curation are required to address idiosyncrasies of source taxonomies, variation in data quality, and conflicting taxonomic opinions. To achieve a harmonized taxonomic view of EOL content, names from data sources are mapped to a dynamic reference hierarchy ([see current version here](<p></p>https://opendata.eol.org/dataset/tram-807-808-809-810-dh-v1-1/resource/00adb47b-57ed-4f6b-8f66-83bfdb5120e8)) using an algorithm that leverages canonical name strings, hierarchical information (ancestry, descendants), taxonomic ranks, synonym data, and author strings. Names that cannot be associated with a reference taxon are still accessible, but their unmapped status excludes them and any associated content from certain core EOL functions. For more information about the EOL taxonomy, see [EOL Dynamic Hierarchy](<p></p>https://eol.org/docs/eol-dynamic-hierarchy)
Test data sets: Amoebozoa Test
The Encyclopedia of Life (EOL, eol.org) aggregates biodiversity information from more than 400 sources and provides access to the data through taxon pages, visual query and application programming interfaces. Scientific names are essential elements of the data integration infrastructure, but their shortcomings as key identifiers are well documented (Patterson et al., 2016). Complex automated workflows and continuous manual curation are required to address idiosyncrasies of source taxonomies, variation in data quality, and conflicting taxonomic opinions. To achieve a harmonized taxonomic view of EOL content, names from data sources are mapped to a dynamic reference hierarchy ([see current version here](<p></p>https://opendata.eol.org/dataset/tram-807-808-809-810-dh-v1-1/resource/00adb47b-57ed-4f6b-8f66-83bfdb5120e8)) using an algorithm that leverages canonical name strings, hierarchical information (ancestry, descendants), taxonomic ranks, synonym data, and author strings. Names that cannot be associated with a reference taxon are still accessible, but their unmapped status excludes them and any associated content from certain core EOL functions. For more information about the EOL taxonomy, see [EOL Dynamic Hierarchy](<p></p>https://eol.org/docs/eol-dynamic-hierarchy)
Integration of data sets from different sources for modeling gender violence and perception of insecurity
<p>The dataset is composed of three distinct files which aggregate processed data derived from open datasets of three cities: Dublin, San Francisco, and Valencia. The data has been mapped to a grid of 25m² for Valencia and 50m² for Dublin and San Francisco. The respective files are named DATA_ES_VLC.csv, DATA_IE_DUB.csv, and DATA_US_SFO.csv. Additionally, there is a dataset for tweets named DATA_TWT.csv, which contains tweets collected through web scraping and analysed using natural language processing (NLP) algorithms and neural networks. The aim is to identify and classify tweets that discuss gender-based violence in the city of Valencia. Another file, MAP_ES_VLC.csv, includes points collected during various mapathons conducted by the Polytechnic University of Valencia campus for a science project aimed at identifying potentially insecure locations.</p>
Ultra-high Resolution Land Use Data Set of Typical Villages in Northeastern Tibetan Plateau
<p>This dataset was collected by a research team during a field investigation in the Hehuang Valley of Qinghai Province from July to August 2022. Using the DJI Mavic2pro equipped with a Hasselblad L1D-20c camera, 55 typical villages were selected in the Hehuang Valley and over 4600 aerial photographs were obtained using drone photogrammetry technology as raw data. Using AgisfphotoScan 1.25 software to synthesize orthophoto images with a spatial resolution of 0.05m. The vector data of human settlement boundaries in villages was extracted through visual interpretation. Based on the object-oriented human-machine interaction interpretation method, 55 typical village land use datasets in 2022 were obtained (including forests, grasslands, forest land, cultivated land, water bodies, roads, unused land, and building land, totaling 8 categories). By establishing 1050 sample points and using confusion matrix analysis, it was found that the overall accuracy of the dataset was 96.86%, with a Kappa coefficient of 0.95. It can accurately reflect the spatial form, land use composition, and surrounding environment of typical villages. Aerial photographs all have longitude, latitude, and altitude information, providing ultra-high resolution data sources for village spatial structure analysis, land use mapping, and analysis work, effectively assisting in the improvement of human housing and rural revitalization strategies.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.