KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I have a packed vector of four 64-bit floating-point values. I would like to get the sum of the vector's elements. With SSE (and using 32-bit floats) I could just do the following: v_sum = _mm_hadd_ps(v_sum, v_sum); v_sum = _mm_hadd_ps(v_sum, v_sum); Unfortunately, even though AVX features a _mm256_hadd_pd instruction, it differs in the result from the SSE version. I believe this is due to the fact that most AVX instructions work as SSE instructions for each low and high 128-bits separately, without ever crossing the 128-bit boundary. Ideally, the solution I am looking for should follow these guidelines: 1) only use AVX/AVX2 instructions. (no SSE) 2) do it in no more than 2-3 instructions. However, any efficient/elegant way to do it (even without following the above guidelines) is always well accepted. Thanks a lot for any help. -Luigi Castelli
Tags (comma-separated)
Save Edits
Cancel