Master the fundamental concepts of gpu architecture through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksSIMD (Single Instruction, Multiple Data) is a processor feature where one instruction operates on a vector of data simultaneously. SSE, AVX, and NEON are SIMD instruction sets. The programmer or compiler explicitly chooses vector width and uses blend/mask instructions to handle per-lane differences.
For example, an AVX-256 add operates on eight 32-bit floats in one instruction. The programmer controls which lanes are active.
SIMT (Single Instruction, Multiple Thread) is how NVIDIA GPUs execute. One instruction broadcasts to many execution units, but you write scalar thread code. Each thread has its own registers and can branch independently. The hardware masks off inactive threads.
Consider a warp of 32 threads where thread_id % 4 == 0 takes one path and the rest take another:
cLoading…
In pure SIMD, a skilled programmer might use blend instructions to handle both paths in one pass. In SIMT, divergence always serializes. AMD wavefronts of 64 threads follow the same rules.
You will simulate SIMT execution with boolean execution masks, show two-phase divergence for an if/else, and model early-exit loops where inactive threads stay masked until all threads finish. This task requires you to print masks per phase and calculate the divergence penalty. Understanding SIMT vs SIMD explains why GPU kernels punish data-dependent branches that CPU vector code can sometimes absorb.
Simulate how a GPU warp executes a divergent program under SIMT. Every lane has its own register r, but the warp has a single instruction stream and an active mask:
if splits the mask and runs the then-side, then the else-side;while keeps looping as long as any lane still wants to, with finished lanes masked off;end.Print every executed statement with the mask it ran under, then each lane's final r.
# starts a comment.
cLoading…
Operands are tid, r or an integer. Statements are echoed with single spaces between their tokens.
[...]: one character per lane, lane 0 first, 1 for active and . for inactive.r from its own values. If any active lane divides or takes % by zero, print the statement, then error: division by zero in an active lane, and stop the run.[mask] if … -> then [T] else [E], with (diverged) when both are non-empty.
else and E is non-empty, print [E] else and run the else-body with E.[mask] end if (reconverged).[mask] while … -> [M]. Stop when M is empty. Otherwise run the body with M.
[mask] end while (reconverged after K iteration(s)).error: loop ran more than 20 iterations and stop the run.step limit reached.r: v0 v1 ….cLoading…
Errors (a program with errors can't run until clear: program has errors):
line N: width 1-32line N: bad assignment, line N: bad conditionline N: else without if, line N: end without if/whileline N: cannot parse "TEXT"run: missing endInput:
cLoading…
Output:
cLoading…
end is exactly the mask before the if or while, as with a hardware reconvergence stack.Hidden tests cover a loop whose lanes exit at different iterations, a uniform branch (empty else-mask), nested divergence, division by zero in an active lane, a runaway loop, and parse errors.