Master the fundamental concepts of profiling & measurement through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksModern CPUs expose performance-monitoring units (PMUs) that count cycles, instructions retired, last-level cache misses, and branch mispredictions. Intel VTune, perf, and likwid all read the same counters. A loop can look innocent in source yet retire few instructions per cycle because cache misses dominate.
Cycles measure wall time on CPU. cache-misses at LLC level often explain memory-bound loops. branch-misses spike when data-dependent branches defeat the predictor.
bashLoading…
perf list shows events your CPU actually supportsperf annotateKeep the relevant documentation open while you implement. When your output disagrees with the reference, trace one failing case by hand before changing random lines.
You will run perf stat on two versions of the same algorithm and compare cycles, cache misses, and branch misses. This exercise requires you to connect counter deltas to specific code changes, not just report raw numbers.
Raw hardware counters from perf stat -e ... mean little on their own. Engineers turn them into ratios: IPC and CPI, cache and branch miss rates, misses per thousand instructions (MPKI), and, when the right events are available, Intel's top-down level-1 breakdown of pipeline slots. Build that analysis step: take recorded counter sets and produce the derived metrics and a verdict.
cLoading…
Counter keys: cycles instructions cache-references cache-misses branches branch-misses L1-dcache-loads L1-dcache-load-misses uops_issued uops_retired idq_uops_not_delivered recovery_cycles.
n/a is printed for a zero denominator. need cycles and instructions.branch mispredictions are costly (>= 5% missed);likely memory bound (IPC < 1, >= 10 cache MPKI);running efficiently (IPC >= 2.5);no single obvious bottleneck.run NAME: bad counter TEXT, report NAME: no such run, and compare: need two runs with cycles.cLoading…
Input:
cLoading…
Output:
cLoading…
Hidden tests cover a branch-heavy run, a frontend-bound top-down profile, a run missing cache counters, L1D metrics, zero denominators, runs assembled over several lines, and malformed counters.