Master the fundamental concepts of simd & vectorization through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksA fair SIMD vs scalar shootout fixes clock, array size, and correctness checks. Report speedup distributions across multiple sizes (L1-fit, L2-fit, RAM) because memory hierarchy changes the winner.
cLoading…
perf stat)For example, AVX2 might show 8x speedup in L1 but only 2x in RAM when DRAM bandwidth caps both paths.
For this exercise, you will run a suite of kernels (add, dot, memset) across scalar, SSE2, and AVX2 implementations. This task asks you to plot speedup vs size and explain where SIMD stops helping, closing the vectorization subtrack with measured evidence.
Keep the relevant datasheet, ISA manual, or architecture textbook chapter open while you implement. When your output disagrees with the reference trace on the same program, the bug is usually a mis-decoded opcode, a stale register read, or a flag bit left unchanged after arithmetic.
For this exercise, you will use those habits while implementing the requirement in the starter code. Microarchitectural product names change across CPU generations, but the control ideas (fetch, bypass, cache lines, vector lanes) stay stable enough to debug from first principles.
Pull the SIMD track together into a roofline-style benchmark model. Each kernel has an arithmetic cost and a memory cost per element. A vector unit divides the arithmetic by its lane count, but the bytes still have to come from whichever level of the memory hierarchy holds the working set. Predict each kernel's cycles for scalar, SSE2 and AVX2 code at several sizes, and explain when SIMD wins and when it hits the bandwidth wall.
cLoading…
N · BYTES. It is served by the first level whose capacity is at least that, else by memory, at bandwidth BW.compute = N · OPS / LANES and memory = N · BYTES / BW cycles. The predicted time is the larger of the two, and that is the bound (compute wins ties).isa line is the baseline.cLoading…
Cycles have one decimal and speedups two.
Input:
cLoading…
Output:
cLoading…
predict(kernel, isa, n) function returning cycles and bound; the report loops over the ISAs in input order.run NAME: unknown kernel.Hidden tests cover a fourth, wider ISA, a working set exactly at a level's capacity, an L2-resident run where SSE2 helps but AVX2 doesn't, and an unknown kernel.