Master the fundamental concepts of profiling & measurement through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksGoogle Benchmark, LLVM's llvm-exegesis, and Julia's @benchmark all share the same skeleton: warmup, many timed iterations, then aggregate statistics. Without warmup, first-touch page faults and cold icache skew the first samples. Without variance reporting, a lucky 5% speedup from turbo noise looks like a win.
Separate setup from the timed body. Loop enough iterations that timer overhead is below 1% of measured time. Track min, median, and standard deviation.
cLoading…
-O2 or -O3 unless you explicitly test debug behaviorvolatile or blackhole sinksKeep the relevant documentation open while you implement. When your output disagrees with the reference, trace one failing case by hand before changing random lines.
You will build a microbenchmark harness with configurable warmup and iteration counts plus mean/min output. This exercise requires honest statistics so you can tell signal from measurement noise.
The statistics are the heart of a microbenchmark framework (Google Benchmark, Criterion, JMH). You discard warmup iterations (cold caches, frequency ramp-up), report robust statistics (median and p90, not just the mean), remove outliers caused by interrupts, flag unstable runs by their coefficient of variation, and refuse to call a difference real when the ranges overlap. Implement that engine over recorded per-iteration timings.
cLoading…
ceil(P * n / 100).sum((x − mean)^2) / (n − 1) (0 for one sample).[Q1 − 3·IQR/2, Q3 + 3·IQR/2] (integer division), where IQR = Q3 − Q1. Recompute mean and sd without them. (unstable: CV above 5%) if CV > 5.0%.warmup: 0..100, analyze NAME: no samples, not enough samples after warmup, and compare: analyze both benchmarks first.cLoading…
The overlap case reads (ranges overlap: the difference may be noise). The warmup count shown is capped at the number of samples.
Input:
cLoading…
Output:
cLoading…
Hidden tests cover an even sample count (median of two), a noisy run flagged as unstable, overlapping ranges, a warmup that swallows every sample, a single sample, and comparing before analysing.