This dataset is a new large-scale high-resolution multi label remote sensing dataset developed by "MLRSNet" for semantic scene understanding. This dataset contains 109161 high-resolution remote sensing images labeled as 46 categories, with sample images ranging from 1500 to 3000 for each category. These images have a fixed size of 256 × 256 pixels and different pixel resolutions. In addition, each image in the dataset is labeled with multiple of the 60 predefined category labels, with each image associated with a label count ranging from 1 to 13. In addition, the construction process of the MLRSNet dataset was elaborated in detail, and the performance of various multi label based deep learning methods in image classification and image retrieval tasks was evaluated. The experimental results show that multi label based deep learning methods can achieve better performance in image classification and image retrieval tasks. This dataset is currently the largest high-resolution multi label remote sensing dataset, containing the richest multi label information. And this dataset has high intra class diversity, which can provide better data resources for the evaluation and development of numerous methods in the field of semantic scene understanding.
| collect place | Global |
|---|---|
| data size | 1.2 GiB |
| data format | jpg,csv |
| Data spatial resolution (/ M) | 10-0.1米 |
MLRSNet consists of 109161 labeled RGB images from around the world, divided into 46 categories: airplanes, airports, bare ground, baseball fields, basketball courts, beaches, bridges, jungles, clouds, commercial areas, densely populated residential areas, deserts, eroded farmland, farmland, forests, highways, golf courses, and ground athletics fields. Ports and harbors, industrial areas, intersections, islands, lakes, grasslands, mobile home parks, mountains, overpasses, parks, parking lots, park avenues, railways, train stations, rivers, roundabouts, shipyards, snow capped mountains, sparse residential areas, sports fields, storage tanks, swimming pools, tennis courts, terraces, transmission towers, vegetable greenhouses, wetlands, and wind turbines. The number of sample images varies greatly depending on different categories, ranging from 1500 to 3000. In addition, each image in the dataset is assigned several of the 60 predefined class labels, and the number of labels associated with each image varies from 1 to 13.
The construction of MLRSNet mainly consists of three processes: scene sample collection, database quality control, and improvement of database sample diversity.
Compared with existing remote sensing image datasets, MLRSNet has the following significant characteristics:
1. Hierarchical structure: MLRSNet consists of three primary categories, such as land use and land cover (e.g. commercial areas, farmland, forests, industrial areas, mountains), natural objects and landforms (e.g. beaches, clouds, islands, lakes, rivers, shrubs), and artificial objects and landforms (e.g. airplanes, airports, bridges, highways, overpasses), with 46 secondary categories and 60 tertiary labels.
2. Multi label: Each image in the MLRSNet dataset corresponds to one or more labels, as remote sensing images typically contain multiple non mutually exclusive object categories. Multiple experiments have shown that in image classification or retrieval tasks, multi label datasets often achieve better performance than single label datasets.
3. Large scale: MLRSNet contains a large number of high-resolution, multi label remote sensing scene images. This dataset contains 109161 high-resolution remote sensing images, labeled as 46 categories, with sample images ranging from 1500 to 3000 for each category, all larger than most other listed datasets. MLRSNet is a large-scale high-resolution remote sensing dataset collected for scene image recognition, which can cover a wider range of satellite or aerial images. It aims to serve as an alternative solution to promote the development of scene image recognition methods, especially deep learning methods that require a large amount of annotated training data.
4. Diversity: To enhance the generalization ability of the dataset, we attempted to describe the features of MLRSNet based on dimensions such as geographic distribution, seasonal distribution, weather conditions, perspective, acquisition time, and image resolution. Specifically, there were significant differences in spatial resolution, perspective, object pose, lighting, background, and occlusion.
This work is licensed under
CC BY 4.0 (Creative Commons Attribution 4.0 International License).
| # | title | file size |
|---|---|---|
| 1 | MLRSNet:用于语义场景理解的多标签高空间分辨率遥感数据集.zip | 1.2 GiB |
Multi label image dataset semantic scene understanding convolutional neural network (CNN) image classification image retrieval
-
-
©Copyright 2005-. Northwest Institute of Eco-Environment and Resources, CAS.
Donggang West Road 320, Lanzhou, Gansu, China (730000)

