KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
Consider the following two snippets: #define ALIGN_BYTES 32 #define ASSUME_ALIGNED(x) x = __builtin_assume_aligned(x, ALIGN_BYTES) void fn0(const float *restrict a0, const float *restrict a1, float *restrict b, int n) { ASSUME_ALIGNED(a0); ASSUME_ALIGNED(a1); ASSUME_ALIGNED(b); for (int i = 0; i < n; ++i) b[i] = a0[i] + a1[i]; } void fn1(const float *restrict *restrict a, float *restrict b, int n) { ASSUME_ALIGNED(a[0]); ASSUME_ALIGNED(a[1]); ASSUME_ALIGNED(b); for (int i = 0; i < n; ++i) b[i] = a[0][i] + a[1][i]; } When I compile the function as gcc-4.7.2 -Ofast -march=native -std=c99 -ftree-vectorizer-verbose=5 -S test.c -Wall I find that GCC inserts aliasing checks for the second function. How can I prevent this such that the resulting assembly for fn1 is the same as that for fn0 ? (When the number of parameters increases from three to, say, 30 the argument-passing approach ( fn0 ) becomes cumbersome and the number of aliasing checks in the fn1 approach becomes ridiculous .) Assembly (x86-64, AVX capable chip); aliasing cruft at .LFB10 fn0: .LFB9: .cfi_startproc testl %ecx, %ecx jle .L1 movl %ecx, %r10d shrl $3, %r10d leal 0(,%r10,8), %r9d testl %r9d, %r9d je .L8 cmpl $7, %ecx jbe .L8 xorl %eax, %eax xorl %r8d, %r8d .p2align 4,,10 .p2align 3 .L4: vmovaps (%rsi,%rax), %ymm0 addl $1, %r8d vaddps (%rdi,%rax), %ymm0, %ymm0 vmovaps %ymm0, (%rdx,%rax) addq $32, %rax cmpl %r8d, %r10d ja .L4 cmpl %r9d, %ecx je .L1 .L3: movslq %r9d, %rax salq $2, %rax addq %rax, %rdi addq %rax, %rsi addq %rax, %rdx xorl %eax, %eax .p2align 4,,10 .p2align 3 .L6: vmovss (%rsi,%rax,4), %xmm0 vaddss (%rdi,%rax,4), %xmm
Tags (comma-separated)
Save Edits
Cancel