Long story short: I'm developing a computing-intensive image processing application in C++. It needs to calculate many variants of image warps on small blocks of pixels extracted from larger images. The program doesn't run as fast as I would like. Profiling (OProfile) showed the warping/interpolation function consumes more than 70% of the CPU time so it seemed obvious to try and optimize it.

I was using the OpenCV image processing library for the task until now:

// some parameters for the image warps (position, stretch, skew)
struct WarpParams;

void Image::get(const WarpParams &params)
{
    // fills matrices mapX_ and mapY_ with x and y coordinates of points to be
    // inteprolated.
    updateCoordMaps(params); 
    // perform interpolation to obtain pixels at point locations 
    // P(mapX_[i], mapY_[i]) based on the original data and put the 
    // result in pixels_. Use bicubic inteprolation.
    cv::remap(image_->data(), pixels_, mapX_, mapY_, CV_INTER_CUBIC);
}

I wrote my own interpolation function and put it in a test harness to ensure correctness while I experiment and to benchmark it in relation to the old one.

My function ran very slow which was to be expected. Generally, the idea is to:

  1. Iterate over the mapX_, mapY_ coordinate maps, extract (real-valued) coordinates of the next pixel to be interpolated;
  2. Retrieve a 4x4 neighborough of pixel values (integer coordinates) from the original image surrounding the interpolated pixel;
  3. Calculate the coefficients of the convolution kernel for each of these 16 pixels;
  4. Calculate the value of the interpolated pixel as a linear combination of the 16 pixel values and the kernel coefficients.

The old function timed at 25us on my Wolfdale Core2 Duo. The new one took 587us (!). I eagerly put my wizard hat on and started hacking the code. I managed to remove all branches, omit some duplicating calculations

Edit
Report