Semantic Segmentation of Aerial Images Using Deep Learning
Extracting building footprints from aerial imagery is a pixel-wise segmentation problem. This article builds a U-Net with a pre-trained VGG-16 encoder to segment buildings in the Inria Aerial Image Labeling Dataset, which covers urban areas in Europe and the United States.
- Cropping 5000×5000 aerial tiles into 512×512 patches for training
- Transfer learning from ImageNet weights, plus augmentation with rotation, zoom and shifts
- Training with Adam and binary cross-entropy, and how the model performs as the dataset grows