How is rotation augmentation implemented efficiently with the overlap-tile strategy in U-Net?
How is rotation augmentation implemented efficiently with the overlap-tile strategy in U-Net?
Loading saved threads...
Juan Carlos Ramírez Tinoco · External communityPost link
External question — Cross Validated Stack Exchange
Author: Juan Carlos Ramírez Tinoco
Original post: https://stats.stackexchange.com/questions/675426
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
When working on image segmentation with Convolutional Neural Networks with high-resolution images, I know there is a trade-off between using full-resolution images and extracting smaller patches to increase batch size during training.
Papers like U-Net (Ronneberger et al., 2015) explain mainly the network architecture, but I wonder if there is work as to specific implementation details regarding data augmentation on these patches.
For instance, the U-Net paper describes the overlap-tile strategy, it states that missing context near the edges of the image is filled using mirroring. However, it does not detail what happens when spatial augmentations, like rotations, are applied during training.
When rotating a square patch by an arbitrary angle, the corners of the new bounding box will fall outside the original pixel data...
My question is: What is the standard practice to handle this in modern pipelines?
From my understanding, there are two potential strategies:
An Oversized Patch: Extract an even larger patch, apply the rotation, and then center-crop the result to the target input size
Whole-Image Rotation: Rotate the entire massive original image first, and then extract the patch (which seems computationally)
Is there a defined standard on this or it just comes down to experimentation?
Quote
Report
Post Reply
Checking account access…