22
What is the advised way of dealing with dynamically-sized datasets in cuda?
Is it a case of 'set the block and grid sizes based on the problem set' or is it worthwhile to assign block dimensions as factors of 2 and have some in-kernel logic to deal with the over-spill?
I can see how this probably matters alot for the block dimensions, but how much does this matter to the grid dimensions? As I understand it, the actual hardware constraints stop at the block level (i.e blocks assigned to SM's that have a set number of SP's, and so can handle a particular warp size).
I've perused Kirk's 'Programming Massively Parallel Processors' but it doesn't really touch on this area.