The Thrust library can be used to sort data. The call might look like this (with a keys and a values vector):

thrust::sort_by_key(d_keys.begin(), d_keys.end(), d_values.begin());

called on the CPU, with d_keys and d_values being in the CPU memory; and the bulk of the execution happens on the GPU.

However, my data is already on the GPU? How can I use the Thrust library to perform efficient sorting directly on the GPU, i.e., to call the sort_by_key function from a kernel?

Also, my data consists of keys that are either unsigned long long int or unsigned int and data that is always unsigned int. How should I make the thrust call for these types?

Edit
Report