Consider the following program:

for i=1 to 10000000 do
  z <- z*z + c

where z and c are complex numbers.

What are efficient x86 assembler implementations of this program using x87 vs SSE and single vs double precision arithmetic?

EDIT I know I can write this in another language and trust the compiler to generate optimal machine code for me but I am doing this to learn how to write optimal x86 assembler myself. I have already looked at the code generated by gcc -O2 and my guess is that there is a lot of room for improvement but I am not adept enough to write optimal x86 assembler by hand myself so I am asking for help here.

Edit
Report