KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Questions tagged
sse
28 questions
Follow Tag
Sort by:
Newest
Votes
Active
9
votes
1
answers
SSE reduction of float vector
c++
sum
sse
simd
reduction
asked 2013-07-20T10:16:49.817
30
votes
1
answers
Why is my hand-tuned, SSE-enabled code so slow?
c++
optimization
opencv
sse
asked 2013-03-30T22:41:37.450
8
votes
1
answers
How to store lower or higher values from AVX/AVX2(YMM) register to memory like the SSE movlps/movhps does?
x86
sse
simd
avx
avx2
asked 2013-01-30T08:23:14.873
18
votes
1
answers
AVX VMOVDQA slower than two SSE MOVDQA?
assembly
sse
bignum
arbitrary-precision
avx
asked 2012-12-20T15:37:04.703
34
votes
4
answers
print a __m128i variable
c
assembly
sse
simd
intrinsics
asked 2012-11-06T18:34:33.680
9
votes
0
answers
attempt to convert SSE2 Fast Corner score code to ARM Neon
arm
sse
neon
computer-vision
asked 2012-08-07T22:08:33.923
18
votes
5
answers
Is it possible to cast floats directly to __m128 if they are 16 byte aligned?
c++
c
alignment
sse
intrinsics
asked 2012-08-01T12:57:14.093
10
votes
4
answers
How to optimize "u[0]*v[0] + u[2]*v[2]" code line with SSE or GLSL
c++
c
optimization
sse
glm-math
asked 2012-06-20T17:37:50.030
12
votes
6
answers
How to allocate 16byte memory aligned data
c
memory
sse
icc
asked 2012-06-18T13:59:59.910
12
votes
2
answers
Is it okay to mix legacy SSE encoded instructions and VEX encoded ones in the same code path?
assembly
x86
sse
avx
asked 2012-06-02T21:21:41.913
22
votes
5
answers
SIMD prefix sum on Intel cpu
c++
sse
simd
prefix-sum
asked 2012-05-14T16:44:36.573
9
votes
3
answers
Efficient complex arithmetic in x86 assembly for a Mandelbrot loop
c
assembly
x86
sse
complex-numbers
asked 2012-04-26T08:39:20.163
10
votes
2
answers
Check XMM register for all zeroes
c++
sse
simd
intrinsics
asked 2012-04-16T14:08:15.900
18
votes
3
answers
How to dump all the XMM registers in gdb?
x86
gdb
simd
sse
cpu-registers
asked 2012-03-30T19:52:43.650
23
votes
3
answers
Fastest way to do horizontal vector sum with AVX instructions
x86
sse
simd
avx
vector-processing
asked 2012-03-19T18:11:59.497
45
votes
4
answers
Using AVX intrinsics instead of SSE does not improve speed -- why?
c++
performance
gcc
sse
avx
asked 2012-01-19T10:47:08.820
60
votes
2
answers
Using AVX CPU instructions: Poor performance without "/arch:AVX"
c++
performance
visual-studio-2010
sse
avx
asked 2011-10-20T17:40:25.700
13
votes
8
answers
Speed up float 5x5 matrix * vector multiplication with SSE
c++
vectorization
matrix-multiplication
sse
simd
asked 2011-07-07T22:04:27.080
9
votes
2
answers
Accepted XX:UseSSE values for Java JVM?
java
jvm
sse
asked 2011-04-05T02:22:51.060
17
votes
5
answers
Does rewriting memcpy/memcmp/... with SIMD instructions make sense?
performance
sse
simd
asked 2011-03-16T05:21:57.310
18
votes
3
answers
Why does does SSE set (_mm_set_ps) reverse the order of arguments
c++
c
simd
sse
intrinsics
asked 2011-03-08T20:30:15.977
19
votes
1
answers
Do I get a performance penalty when mixing SSE integer/float SIMD instructions
c
assembly
sse
simd
intrinsics
asked 2011-02-14T19:28:05.790
14
votes
8
answers
Compute the absolute difference between unsigned integers using SSE
c++
unsigned
sse
asked 2010-08-01T04:52:52.033
13
votes
6
answers
What is the fastest way to test if a double number is integer (in modern intel X86 processors)
c
optimization
assembly
x86
sse
asked 2009-12-22T04:09:05.633
10
votes
8
answers
SIMD programming languages
programming-languages
sse
simd
ispc
asked 2009-09-13T12:50:45.267
56
votes
5
answers
SSE SSE2 and SSE3 for GNU C++
c++
optimization
simd
sse
sse2
asked 2009-03-19T07:32:22.857
12
votes
3
answers
How to get GCC to use more than two SIMD registers when using intrinsics?
gcc
assembly
x86
sse
simd
asked 2008-09-23T22:49:22.830
173
votes
4
answers
What is the meaning of "non temporal" memory accesses in x86
x86
sse
assembly
asked 2008-08-31T20:18:34.113