Alex Rivera | Logout

How to parallelize list-comprehension calculations in Python?

Asked 2011-03-08T17:57:12.620
68

Both list comprehensions and map-calculations should -- at least in theory -- be relatively easy to parallelize: each calculation inside a list-comprehension could be done independent of the calculation of all the other elements. For example in the expression

[ x*x for x in range(1000) ]

each x*x-Calculation could (at least in theory) be done in parallel.

My question is: Is there any Python-Module / Python-Implementation / Python Programming-Trick to parallelize a list-comprehension calculation (in order to use all 16 / 32 / ... cores or distribute the calculation over a Computer-Grid or over a Cloud)?

Edit
Report

2 Answers

4

No, because list comprehension itself is a sort of a C-optimized macro. If you pull it out and parallelize it, then it's not a list comprehension, it's just a good old fashioned MapReduce.

But you can easily parallelize your example. Here's a good tutorial on using MapReduce with Python's parallelization library:

http://mikecvet.wordpress.com/2010/07/02/parallel-mapreduce-in-python/

answered 2011-03-08T18:43:09.097
1

There is a comprehensive list of parallel packages for Python here:

http://wiki.python.org/moin/ParallelProcessing

I'm not sure if any handle the splitting of a list comprehension construct directly, but it should be trivial to formulate the same problem in a non-list comprehension way that can be easily forked to a number of different processors. I'm not familiar with cloud computing parallelization, but I've had some success with mpi4py on multi-core machines and over clusters. The biggest issue that you'll have to think about is whether the communication overhead is going to kill any gains you get from parallelizing the problem.

Edit: The following might also be of interest:

http://www.mblondel.org/journal/2009/11/27/easy-parallelization-with-data-decomposition/

answered 2011-03-08T18:03:23.413

Your Answer