Master the fundamental concepts of memory hierarchy through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksMulti-socket servers attach local DRAM to each CPU. Accessing another socket's memory crosses an interconnect (Intel UPI, AMD Infinity Fabric) with higher latency and limited bandwidth. The OS and allocator should prefer local nodes.
bashLoading…
For example, allocating on node 0 then letting threads on node 1 hammer the buffer can halve bandwidth versus node-local allocation.
For this exercise, you will run the same bandwidth microbenchmark under numactl policies and compare nodes. This task asks you to quantify why PostgreSQL, Redis, and HPC jobs pin threads near their memory.
Keep the relevant datasheet, ISA manual, or architecture textbook chapter open while you implement. When your output disagrees with the reference trace on the same program, the bug is usually a mis-decoded opcode, a stale register read, or a flag bit left unchanged after arithmetic.
For this exercise, you will use those habits while implementing the requirement in the starter code. Microarchitectural product names change across CPU generations, but the control ideas (fetch, bypass, cache lines, vector lanes) stay stable enough to debug from first principles.
Simulate memory placement on a NUMA machine. Each socket (node) has its own memory. Local accesses are fast and remote ones cost extra, and the operating system chooses where each page lives. Model Linux's allocation policies (first-touch, interleave and bind) and thread pinning. Then show the classic bug, where one thread initialises the data and so places every page on its own node, together with its fixes.
cLoading…
The policy in force when a buffer is allocated is the one that applies to that buffer's pages. The default is first-touch.
i of the buffer goes to node i mod N.K.Once placed, a page never moves.
For every read/write:
cLoading…
placed counts pages placed by this operation. Each page access costs L if the page is on the thread's node, else R, and avg is their mean. show BUF prints BUF: n n n ... (each page's node, - if unplaced). At the end:
cLoading…
Percentages and averages have one decimal.
Input:
cLoading…
Output:
cLoading…
Hidden tests cover parallel first-touch initialisation (each worker initialises its own half), the interleave policy with three nodes, bind, and a thread that is migrated to another node after its data was placed.