8
To elaborate on Robert's answer, here's an example of how you could use streams to make your two instances of Kernel1 run concurrently:
cudaStream_t stream1; cudaStreamCreate(&stream1);
cudaStream_t stream2; cudaStreamCreate(&stream2);
Kernel1<<gridDim, blockDim, 0, stream1>>(dst1, param1);
Kernel1<<gridDim, blockDim, 0, stream2>>(dst2, param2);
A few more notes about concurrent execution with streams:
Kernel1<<<g, b>>>(), and then launch a kernel with a specific stream Kernel2<<<g, b, 0, stream>>>(), then Kernel2 will wait for Kernel1 to finish. Kernel1<<<g, b>>>()), Nvidia calls this "using the NULL stream."cudaEvents, your work can sometimes get serialized even if you distribute the kernels over several streams.