CNN with fixed batch size - repeat to fill or reduce batch size?

CNN with fixed batch size - repeat to fill or reduce batch size?

Manage alerts

Loading saved threads...

User1291 · External communityPost link
External question — Cross Validated Stack Exchange Author: User1291 Original post: https://stats.stackexchange.com/questions/279642 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I'm setting up a CNN that can handle differently-sized examples so long as all examples within the same batch are of the same size. As a trade-off for this flexibility, I need to fix the batch size. (Only one dimension of the tensor can be dynamic). In the current implementation, the training examples are split into length-keyed buckets. One iteration then consists of randomly sampling BATCH_SIZE many examples from each bucket and feeding these batches - in turn - through the network. I am unsure, however, what I should do if I don't have enough samples of a certain length to fill that batch with distinct examples. Say, I may have 55'000 examples but only 7 of length 2 when the batch size is 50. Should I drop lengths for which I don't have "enough" examples? Should I just keep sampling from however many examples I have until I have a full batch (obviously repeating some examples)? Pre-process the input examples and choose smallest possible batch size (at risk of that turning out to be 1)? Or something else entirely? I'm currently leaning towards "just keep sampling until the batch is full" because we're losing information otherwise. However, I am also worried that this might skew the network somehow. Somehow , I say, because I quite simply don't know and am ill-equipped to make an educated guess. Are these concerns founded?
Quote
Report
Bruno Lubascher · External communityPost link
External answer — Cross Validated Stack Exchange Author: Bruno Lubascher Original post: https://stats.stackexchange.com/a/279644 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Without knowing what kind of data you have, I would suggest that you try zero padding your input data. In this way, all your examples have the same size and your network is more robust to the different sizes of inputs, even to those that you do not have any training examples for. In that case, you would fix your input size to the largest training example you have (or a larger number if you know that your data could come in a larger format). With this, you would be able to vary your batch size.
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: User1291 Source score (net votes, not local likes): 1 Original post: https://stats.stackexchange.com/questions/279642 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I'm setting up a CNN that can handle differently-sized examples so long as all examples within the same batch are of the same size. As a trade-off for this flexibility, I need to fix the batch size. (Only one dimension of the tensor can be dynamic). In the current implementation, the training examples are split into length-keyed buckets. One iteration then consists of randomly sampling BATCH_SIZE many examples from each bucket and feeding these batches - in turn - through the network. I am unsure, however, what I should do if I don't have enough samples of a certain length to fill that batch with distinct examples. Say, I may have 55'000 examples but only 7 of length 2 when the batch size is 50. Should I drop lengths for which I don't have "enough" examples? Should I just keep sampling from however many examples I have until I have a full batch (obviously repeating some examples)? Pre-process the input examples and choose smallest possible batch size (at risk of that turning out to be 1)? Or something else entirely? I'm currently leaning towards "just keep sampling until the batch is full" because we're losing information otherwise. However, I am also worried that this might skew the network somehow. Somehow , I say, because I quite simply don't know and am ill-equipped to make an educated guess. Are these concerns founded?

Cancel quote

Checking account access…