Alex Rivera | Logout

How to use coalesced memory access

Asked 2011-07-03T13:51:24.687
9

I have 'N' threads to perform simultaneously on device which they need M*N float from the global memory. What is the correct way to access the global memory coalesced? In this matter, how the shared memory can help?

Edit
Report

1 Answer

17

Usually, a good coalesced access can be achieved when the neighbouring threads access neighbouring cells in memory. So, if tid holds the index of your thread, then accessing:

  • arr[tid] --- gives perfect coalescence
  • arr[tid+5] --- is almost perfect, probably misaligned
  • arr[tid*4] --- is not that good anymore, because of the gaps
  • arr[random(0..N)] --- horrible!

I am talking from the perspective of a CUDA programmer, but similar rules apply elsewhere as well, even in a simple CPU programming, although the impact is not that big there.


"But I have so many arrays everyone has about 2 or 3 times longer than the number of my threads and using the pattern like "arr[tid*4]" is inevitable. What may be a cure for this?"

If the offset is a multiple of some higher 2-power (e.g. 16*x or 32*x) it is not a problem. So, if you have to process a rather long array in a for-loop, you can do something like this:

for (size_t base=0; i<arraySize; i+=numberOfThreads)
    process(arr[base+threadIndex])

(the above asumes that array size is a multiple of the number of threads)

So, if the number of threads is a multiple of 32, the memory access will be good.

Note again: I am talking from the perspective of a CUDA programmer. For different GPUs/environment you might need less or more threads for perfect memory access coalescence, but similar rules should apply.


Is "32" related to the warp size which access parallel to the global memory?

Although not directly, there is some connection. Global memory is divided into segments of 32, 64 and 128 bytes which are accessed by half-warps. The more segments you access for a given memory-fetch instruction, the longer it goes. You can read more into details in the "CUDA Programming Guide", there is a whole chapter on

answered 2011-07-03T15:09:39.943

Your Answer