Master the fundamental concepts of simd & vectorization through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksAVX2 widens integer and floating SIMD to 256 bits (__m256i). One VADDPS fuses eight 32-bit floats, doubling theoretical throughput versus SSE on the same core (when ports support AVX).
cLoading…
Some Intel CPUs downclock when AVX-heavy code runs (AVX-512 more so). Measure end-to-end time, not just instruction count.
_mm256_load_ps-mavx2 -mfma as neededFor this exercise, you will port your SSE2 kernel to AVX2 and compare GB/s. This task asks you to handle remainder elements scalarly while keeping the hot loop on __m256 operations.
Keep the relevant datasheet, ISA manual, or architecture textbook chapter open while you implement. When your output disagrees with the reference trace on the same program, the bug is usually a mis-decoded opcode, a stale register read, or a flag bit left unchanged after arithmetic.
For this exercise, you will use those habits while implementing the requirement in the starter code. Microarchitectural product names change across CPU generations, but the control ideas (fetch, bypass, cache lines, vector lanes) stay stable enough to debug from first principles.
Upgrade the vector kernel from SSE2 (4 floats per 128-bit register) to AVX2 (8 floats per 256-bit register), and choose the implementation at runtime from the CPU's features, the way libraries dispatch on cpuid. Wider vectors halve the number of loop iterations, but they need 32-byte alignment, so the scalar prologue can be longer. For short arrays the wider path can even lose.
cLoading…
A cpu line applies to the run lines after it.
float arrays: a[i] = 0.5·i − 3, b[i] = (i mod 5) + 0.25, and c[i] = a[i] OP b[i]. With these values every result is exact in single precision.
For a path with W lanes and alignment A bytes: prologue P = ((A − OFFSET mod A) mod A) / 4 (at most N), body V = ⌊(N − P)/W⌋ vectors, epilogue E = N − P − W·V. Its cost is P + V + E operations. The scalar path costs N.
W = 4, A = 16W = 8, A = 32Dispatch: use avx2 if the CPU has it, else sse2, else scalar.
After a cpu line: cpu: FEATURES -> dispatch PATH, listing the features as given (or none). Per run:
cLoading…
Both vector paths are always reported. using names the dispatched path and its cost (for scalar, N ops), and the speedup is N / cost with two decimals. Print the first min(N, 4) results and the checksum with three decimals (every value here is a multiple of 1/8, so they are exact). The checksum is the sum of all c[i], accumulated in a double in index order.
Input:
cLoading…
Output:
cLoading…
cpu line.OFFSET is simulated.Hidden tests cover a CPU with only SSE2, one with neither (scalar fallback), mul, offsets where avx2 has the longer prologue, and a large array where avx2 approaches 8×.