Master the fundamental concepts of gpu architecture through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksA modern GPU can issue hundreds of billions of floating-point operations per second, but global memory bandwidth is typically 1-2 terabytes per second. With thousands of concurrent threads, each thread gets only a thin slice of that bandwidth unless accesses are merged efficiently.
Memory coalescing is the mechanism that combines requests from threads in the same warp into as few cache-line transactions as possible. Coalescing is about transaction count, not bytes transferred.
Consider 32 threads in a warp, each reading a 4-byte integer with 128-byte cache lines:
cLoading…
For example, in pattern (a), threads 0-31 read addresses base+0 through base+124, all within one 128-byte line. In pattern (b), each thread touches a different line despite requesting the same total bytes.
You will implement a transaction counter for coalesced, strided, and misaligned patterns, then compare row-major vs column-major matrix reads. This task requires you to print a summary table with efficiency percentages. The analysis you build here is the same reasoning CUDA kernel authors apply when nvcc warns about uncoalesced global loads.
When a warp executes a load, the hardware merges its 32 threads' addresses into as few memory transactions as possible. This is coalescing: global memory moves whole 32-byte sectors, grouped in 128-byte lines. Shared memory instead has 32 banks of 4-byte words: threads hitting different words in the same bank are serialized (a bank conflict), while threads reading the same word get a broadcast. Compute both costs for common access patterns.
# starts a comment.
cLoading…
byte / 32) and lines (byte / 128) covered by every active thread's byte range. Then:
struct of arrays, same field, where thread t loads the field at t × FIELD_SIZE.cLoading…
thread(s), sector(s), line(s).active: 1-32strided: BASE STRIDE [SIZE 1-16]aos: STRUCT_SIZE FIELD_OFFSET FIELD_SIZE (the field must fit in the struct)gather: SIZE ADDR..., gather: need N addresses, gather: addresses must be >= 0shared: BASE_WORD STRIDE_WORDS, transpose: TILE PAD (TILE 1-64, PAD 0-8)unknown command XInput:
cLoading…
Output:
cLoading…
Hidden tests cover a misaligned base, large strides, 16-byte vector loads, a broadcast (stride 0), a partial warp, a gather, padded and unpadded transposes, and argument errors.