KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I am working on a language that is compiled with LLVM. Just for fun, I wanted to do some microbenchmarks. In one, I run some million sin / cos computations in a loop. In pseudocode, it looks like this: var x: Double = 0.0 for (i <- 0 to 100 000 000) x = sin(x)^2 + cos(x)^2 return x.toInteger If I'm computing sin/cos using LLVM IR inline assembly in the form: %sc = call { double, double } asm "fsincos", "={st(1)},={st},1,~{dirflag},~{fpsr},~{flags}" (double %"res") nounwind this is faster than using fsin and fcos separately instead of fsincos. However, it is slower than if I calling the llvm.sin.f64 and llvm.cos.f64 intrinsics separately, which compile to calls to the C math lib functions, at least with the target settings I'm using (x86_64 with SSE enabled). It seems LLVM inserts some conversions between single/double precision FP -- that might be the culprit. Why is that? Sorry, I'm a relative newbie at assembly: .globl main .align 16, 0x90 .type main,@function main: # @main .cfi_startproc # BB#0: # %loopEntry1 xorps %xmm0, %xmm0 movl $-1, %eax jmp .LBB44_1 .align 16, 0x90 .LBB44_2: # %then4 # in Loop: Header=BB44_1 Depth=1 movss %xmm0, -4(%rsp) flds -4(%rsp) #APP fsincos #NO_APP fstpl -16(%rsp) fstpl -24(%rsp) movsd -16(%rsp), %xmm0 mulsd %xmm0, %xmm0 cvtsd2ss %xmm0, %xmm1 movsd -24(%rsp), %xmm0 mulsd %xmm0, %xmm0 cvtsd2ss %xmm0, %xmm0 addss %xmm1, %xmm0 .LBB44_1: # %loop2 # =>This Inner Loop Header: Depth=1 incl %eax cmpl $99999999, %eax # imm = 0x5F5E0FF jle .LBB44_2 # BB#3:
Tags (comma-separated)
Save Edits
Cancel