KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I'm currently parallelizing program using openmp on a 4-core phenom2. However I noticed that my parallelization does not do anything for the performance. Naturally I assumed I missed something (falsesharing, serialization through locks, ...), however I was unable to find anything like that. Furthermore from the CPU Utilization it seemed like the program was executed on only one core. From what I found sched_getcpu() should give me the Id of the core the thread executing the call is currently scheduled on. So I wrote the following test program: #include <iostream> #include <sstream> #include <omp.h> #include <utmpx.h> #include <random> int main(){ #pragma omp parallel { std::default_random_engine rand; int num = 0; #pragma omp for for(size_t i = 0; i < 1000000000; ++i) num += rand(); auto cpu = sched_getcpu(); std::ostringstream os; os<<"\nThread "<<omp_get_thread_num()<<" on cpu "<<sched_getcpu()<<std::endl; std::cout<<os.str()<<std::flush; std::cout<<num; } } On my machine this gives the following output(the random numbers will vary of course): Thread 2 on cpu 0 num 127392776 Thread 0 on cpu 0 num 1980891664 Thread 3 on cpu 0 num 431821313 Thread 1 on cpu 0 num -1976497224 From this I assume that all threads execute on the same core (the one with id 0). To be more certain I also tried the approach from this answer . The results where the same. Additionally using #pragma omp parallel num_threads(1) didn't make the execution slower (slightly faster in fact), lending credibility to the theory that all threads use the same cpu, however the fact that the cpu is always displayed as 0 makes me kind of suspicious. Additionally I checked GOMP_CPU_AFFINITY</co
Tags (comma-separated)
Save Edits
Cancel