Master the fundamental concepts of simd & vectorization through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksModern Clang and GCC scan loops for SIMD opportunities when -O3 -march=native and #pragma omp simd or __restrict__ prove independence. The compiler emits AVX2/NEON instructions you never wrote in source.
cLoading…
c might overlap a)For example, adding __restrict__ on pointers often unlocks vectorization identical to hand intrinsics.
For this exercise, you will toggle pragmas and compiler flags on the same kernel and dump assembly with objdump -d. This task asks you to document one loop the compiler vectorizes and one it refuses, with the reported reason from optimization remarks.
Keep the relevant datasheet, ISA manual, or architecture textbook chapter open while you implement. When your output disagrees with the reference trace on the same program, the bug is usually a mis-decoded opcode, a stale register read, or a flag bit left unchanged after arithmetic.
For this exercise, you will use those habits while implementing the requirement in the starter code. Microarchitectural product names change across CPU generations, but the control ideas (fetch, bypass, cache lines, vector lanes) stay stable enough to debug from first principles.
Compilers vectorise loops automatically, but only when they can prove it is safe. Write the legality analysis of an auto-vectoriser like GCC's. Given the compiler flags and a one-statement loop body, decide whether the loop can be vectorised for SSE (16-byte vectors). If it can, report the extra machinery needed (masking, reduction, a runtime alias check). If it can't, give the reason. This is the reasoning you need to read -fopt-info-vec-missed output and to know when restrict or -ffast-math will help.
cLoading…
T is char, short, int, float, long or double, the element type of every array. The loop runs over i. restrict means no two arrays overlap. STATEMENT is one of:
X[i] = EXPR: a store; EXPR may use array elements, constants and calls such as sqrtf(b[i]);s += EXPR: a reduction into the scalar s;if (COND) X[i] = EXPR: a conditional store.Array indices are i, i+K, i-K, or anything else, which is indirect, e.g. b[c[i]].
VF (the vectorisation factor) is 16 / sizeof(T): char 16, short 8, int and float 4, long and double 2.
-O3 → not vectorized: the vectorizer runs at -O3.not vectorized: call to NAME (the first one).not vectorized: indirect access X[...] (gather/scatter) (the first such array, in reading order; for a nested index like b[c[i]] that's b).i − d (d ≥ 1) is a loop-carried dependence. If the smallest such d is below VF → not vectorized: dependence distance d < VF V. Reads at i or i + K are always safe.float/double without -ffast-math → not vectorized: T reduction needs -ffast-math (reassociation) (integer reductions are fine). + reduction for a reduction, + masking for an if, + distance d >= VF when a dependence exists but is far enough, and + runtime alias check when the statement stores to an array, reads another array, and the loop isn't restrict (the compiler versions the loop).flags: -O3 -ffast-math for a flags line (or flags: (no -O3) plus -ffast-math if present), and per loop:
cLoading…
Input:
cLoading…
Output:
cLoading…
Hidden tests cover integer reductions, -ffast-math, conditional stores, calls, gathers and scatters, a dependence exactly at VF, forward (i + K) reads, char loops with VF = 16, and a loop compiled without -O3.