Master the fundamental concepts of gpu architecture through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksGPUs expose a multi-level memory hierarchy that programmers often manage explicitly. Understanding each level is essential for writing efficient compute kernels and shaders.
If a kernel uses too many registers, the compiler spills them to local memory backed by VRAM. That spill is orders of magnitude slower than keeping values in registers. This tradeoff is called register pressure.
GPUs have enormous bandwidth but terrible latency. The scheduler hides latency by switching to other ready warps while one warp waits on a global load. For example, a kernel that performs one global read per thread with low occupancy may stall frequently because too few warps are available to cover the ~400 cycle VRAM round trip.
cLoading…
Shared memory requires explicit allocation and copy. L1 cache is automatic but smaller. The key insight: maximize active warps and minimize global traffic through shared memory staging and coalesced access.
You will model latency and bandwidth at each hierarchy level and compute how many warps a kernel needs to hide global memory stalls. This task asks you to print per-level access costs and explain when register spills would hurt occupancy. The numbers you compute here connect directly to the coalescing and occupancy exercises that follow.
Model the GPU memory hierarchy well enough to see why kernels are written the way they are:
Simulate individual global loads through fully-associative LRU caches, sweep over memory with different strides, and estimate how shared-memory tiling cuts the global traffic of a matrix multiply.
# starts a comment.
cLoading…
addr / 128.cLoading…
sweep is truncated. The hit percentage is rounded half up to 1 decimal (0.0% with no loads).access for 1 access and accesses otherwise.latency: LEVEL CYCLES ... (levels reg shared l1 l2 dram)l1: LINES (1-64) (or l2)load: ADDR, sweep: START COUNT STRIDEreg: COUNT (or shared)matmul: N TILE (TILE must divide N, N <= 4096)unknown command XInput:
cLoading…
Output:
cLoading…
Hidden tests cover LRU eviction with tiny caches (L2 catching L1 misses), changed latencies, a sweep that thrashes the caches, a full-tile matmul, and argument errors.