Master the fundamental concepts of profiling & measurement through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksIntel VTune and AMD uProf add source-line mapping, memory access analysis, and microarchitecture-specific metrics like frontend stalls and port utilization. Kernel teams use them when perf events are multiplexed too aggressively or when you need cache-line conflict detail that generic counters miss.
Build with debug symbols (-g). Run the profiler in hotspot or memory-access mode. Drill from function to source line to assembly. Compare before/after optimization runs saved as separate sessions.
bashLoading…
Keep the relevant documentation open while you implement. When your output disagrees with the reference, trace one failing case by hand before changing random lines.
You will document a profiling plan using VTune or uProf on a sample binary and interpret a provided summary table. This exercise requires mapping one hotspot function to a concrete optimization hypothesis.
VTune's Hotspots and Microarchitecture views, and AMD uProf's, are built from event-based sampling. Every N occurrences of an event (the sample-after value, SAV), the PMU records the current instruction pointer. Samples × SAV estimates the event count per function and source line. Ratios of those estimates (CPI, LLC misses per kilo-instruction) show why a hotspot is slow. Build the report from recorded samples.
cLoading…
The default SAVs are cycles 2000003, instructions 2000003, and llc_miss 20011 (VTune uses odd periods, which avoids aliasing with loops).
n/a. (no instruction samples); -> memory bound: improve data locality; -> stalled, but not on LLC misses: check dependencies and branches; -> retiring well: reduce the instruction count (vectorise);lines FUNC: no cycle samples.sav: unknown event or bad period and sample: bad event, line or count.cLoading…
Input:
cLoading…
Output:
cLoading…
Hidden tests cover custom SAVs that change the CPI, a function with no instruction samples, ties in the ranking, hotspots 1, per-line tables with ties, and invalid samples.