Master the fundamental concepts of gpu architecture through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksOccupancy is the ratio of active warps on an SM to the maximum warps the SM can hold. Higher occupancy means more warps are available to hide memory latency, but it is not the only performance metric. A kernel with 100% occupancy can still be slow if every thread does redundant work.
The formula is straightforward:
cLoading…
For example, if an SM supports 48 warps and your kernel places 24 warps on it, occupancy is 50%.
Three resources compete to determine how many thread blocks fit on one SM:
registers_per_thread * threads_per_blockthreads_per_block / 32, capped by max warps per SMIf a kernel uses 128 registers per thread with 512 threads per block, register demand alone may allow only one block on a 64K-register SM. Shared memory can be the bottleneck too: four blocks at 24KB each fill a 96KB pool even when registers are plentiful.
You will read registers_per_thread, shared_memory_per_block, and threads_per_block from stdin and compute occupancy, identifying whether registers, shared memory, or warps is the limiting factor. This task asks you to test three configurations and explain whether higher occupancy would help. The calculator you build mirrors what nvcc --ptxas-options=-v reports when CUDA authors debug register pressure.
Write the core of NVIDIA's occupancy calculator. Given a GPU architecture and a kernel's launch shape (threads per block, registers per thread, shared memory per block), compute how many blocks fit on one SM under each resource limit, the resulting occupancy (active warps / max warps), which resource is the limiter, and what change would fit one more block.
| name | max warps/SM | max blocks/SM | registers/SM | reg unit | shared/SM (B) | shared unit | max shared/block | reserved/block |
|---|---|---|---|---|---|---|---|---|
| sm_75 | 32 | 16 | 65536 | 256 | 65536 | 256 | 65536 | 0 |
| sm_80 | 64 | 32 | 65536 | 256 | 167936 | 128 | 166912 | 1024 |
| sm_86 | 48 | 16 | 65536 | 256 | 102400 | 128 | 101376 | 1024 |
| sm_90 | 64 | 32 | 65536 | 256 | 233472 | 128 | 232448 | 1024 |
All of them allow at most 1024 threads per block and 255 registers per thread. The default is sm_80.
# starts a comment.
cLoading…
ceil(THREADS / 32).floor(max_warps / wpb). by blocks: max_blocks.REGS × 32 rounded up to the reg unit; blocks = floor(floor(65536 / regs_per_warp) / wpb).SHARED + reserved rounded up to the shared unit; blocks = floor(shared_per_SM / bytes). With 0 bytes per block, there is no limit.and.floor(shared_per_SM / (active+1)) rounded down to the shared unit, minus the reserved bytes. Print it only if it is ≥ 0 and below SHARED.cLoading…
incl. N reserved appears only when the architecture reserves memory. unlimited (0 B per block) replaces the block count when no shared memory is used. (the kernel cannot launch) before - limited by.block/blocks everywhere.gpu: sm_75, sm_80, sm_86 or sm_90kernel: THREADS REGS SHARED_BYTESkernel: threads per block must be 1-1024kernel: registers per thread must be 1-255kernel: shared memory per block must be 0-MAX bytesunknown command XInput:
cLoading…
Output:
cLoading…
Hidden tests cover an architecture without reserved shared memory, a 255-register kernel that cannot launch, a shared-memory-limited kernel with its hint, a block-count-limited small kernel, tied limiters, and every input error.