Extracting building footprints from aerial imagery is a pixel-wise segmentation problem. This article builds a U-Net with a pre-trained VGG-16 encoder to segment buildings in the Inria Aerial Image Labeling Dataset, which covers urban areas in Europe and the United States.

  • Cropping 5000×5000 aerial tiles into 512×512 patches for training
  • Transfer learning from ImageNet weights, plus augmentation with rotation, zoom and shifts
  • Training with Adam and binary cross-entropy, and how the model performs as the dataset grows

Read the full article on Medium

Updated: