Alex Rivera | Logout

How Do You Profile & Optimize CUDA Kernels?

Asked 2010-02-05T01:40:16.373
22

I am somewhat familiar with the CUDA visual profiler and the occupancy spreadsheet, although I am probably not leveraging them as well as I could. Profiling & optimizing CUDA code is not like profiling & optimizing code that runs on a CPU. So I am hoping to learn from your experiences about how to get the most out of my code.

There was a post recently looking for the fastest possible code to identify self numbers, and I provided a CUDA implementation. I'm not satisfied that this code is as fast as it can be, but I'm at a loss as to figure out both what the right questions are and what tool I can get the answers from.

How do you identify ways to make your CUDA kernels perform faster?

Edit
Report

1 Answer

2

The CUDA profiler is rather crude and doesn't provide a lot of useful information. The only way to seriously micro-optimize your code (assuming you have already chosen the best possible algorithm) is to have a deep understanding of the GPU architecture, particularly with regard to using shared memory, external memory access patterns, register usage, thread occupancy, warps, etc.

Maybe you could post your kernel code here and get some feedback ?

The nVidia CUDA developer forum forum is also a good place to go for help with this kind of problem.

answered 2010-02-05T08:41:24.967

Your Answer