19
I want a measure of how much of the peak performance my kernel archives.
Say I have a NVIDIA Tesla C1060, which has a peak GFLOPS of 622.08 (~= 240Cores * 1300MHz * 2). Now in my kernel I counted for each thread 16000 flop (4000 x (2 subtraction, 1 multiplication and 1 sqrt)). So when I have 1,000,000 threads I would come up with 16GFLOP. And as the kernel takes 0.1 seconds I would archive 160GFLOPS, which would be a quarter of the peak performance. Now my questions:
- Is this approach correct?
- What about comparisons (
if(a>b) then....)? Do I have to consider them as well? - Can I use the CUDA profiler for easier and more accurate results? I tried the
instructionscounter, but I could not figure out, what the figure means.
sister question: How to calculate the achieved bandwidth of a CUDA kernel