KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I'm constructing a micro-benchmark to measure performance changes as I experiment with the use of SIMD instruction intrinsics in some primitive image processing operations. However, writing useful micro-benchmarks is difficult, so I'd like to first understand (and if possible eliminate) as many sources of variation and error as possible. One factor that I have to account for is the overhead of the measurement code itself. I'm measuring with RDTSC, and I'm using the following code to find the measurement overhead: extern inline unsigned long long __attribute__((always_inline)) rdtsc64() { unsigned int hi, lo; __asm__ __volatile__( "xorl %%eax, %%eax\n\t" "cpuid\n\t" "rdtsc" : "=a"(lo), "=d"(hi) : /* no inputs */ : "rbx", "rcx"); return ((unsigned long long)hi << 32ull) | (unsigned long long)lo; } unsigned int find_rdtsc_overhead() { const int trials = 1000000; std::vector<unsigned long long> times; times.resize(trials, 0.0); for (int i = 0; i < trials; ++i) { unsigned long long t_begin = rdtsc64(); unsigned long long t_end = rdtsc64(); times[i] = (t_end - t_begin); } // print frequencies of cycle counts } When running this code, I get output like this: Frequency of occurrence (for 1000000 trials): 234 cycles (counted 28 times) 243 cycles (counted 875703 times) 252 cycles (counted 124194 times) 261 cycles (counted 37 times) 270 cycles (counted 2 times) 693 cycles (counted 1 times) 1611 cycles (counted 1 times) 1665 cycles (counted 1 times) ... (a bunch of larger times each only seen once) My questions are these: What are the possible causes of the bi-modal distribution of cycle counts generated by the code above? Why does the fastest time (234 cycles) only occur a handful of times—what highly unus
Tags (comma-separated)
Save Edits
Cancel