Master the fundamental concepts of binary formats through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksDWARF is the standard debugging data format used by GCC and Clang on Unix-like systems. It stores information about types, variables, functions, and source line numbers inside sections like .debug_info and .debug_line.
DWARF organizes debug data as a tree of Debugging Information Entries (DIEs):
DW_TAG_compile_unit: Root of a compilation unitDW_TAG_subprogram: A function or subroutineDW_TAG_variable: A variableDW_AT_name: Name of the entityDW_AT_low_pc / DW_AT_high_pc: Function address rangeFor example, scanning .debug_info for DW_TAG_subprogram entries with DW_AT_name reveals every function name even when the symbol table has been stripped.
You will implement a simplified DWARF scanner that extracts function names from a .debug_info section. Full DWARF parsing is complex, but this exercise teaches you what GDB uses to map instruction addresses back to source lines.
Beyond function names, .debug_line maps instruction addresses to source file and line number. GDB uses this mapping when you break main.c:42. The abbreviation table in each compilation unit compresses repeated DIE patterns, which is why raw .debug_info bytes look opaque until you decode the abbrev entries. Even a simplified scanner teaches you what debuggers rely on under the hood.
Scan a simplified .debug_info stream the way a debugger does, and list the functions it describes. It keeps DWARF's core mechanics: tags, ULEB128 variable-length integers, names resolved through a string table, and DW_AT_high_pc in offset form.
Every record (a DIE) has the same layout:
| Field | Size |
|---|---|
| tag | 1 byte |
| name offset into the string table | ULEB128: 7 bits per byte, low bits first; a set high bit means another byte follows, so offsets above 127 take 2+ bytes |
| low_pc | 8 bytes, little-endian |
| high_pc offset | 8 bytes, little-endian; the function ends at low_pc + offset |
Tag 0x11 is DW_TAG_compile_unit, 0x2E is DW_TAG_subprogram, and any other tag (such as 0x34 DW_TAG_variable) is skipped. A name offset past the end of the string table resolves to ?.
cLoading…
then one line per DIE:
CU: main.c"FUNC %-12s low=0x%llX high=0x%llX size=%llu\n": FUNC factorial low=0x401180 high=0x4011C0 size=64SKIP tag=0x34 name_off=45, the tag as 2 uppercase hex digits and the offset in decimalThen Summary: 1 CUs, 3 subprograms, 197 bytes of code (the sum of the subprogram sizes) and, if there was any compile unit, Compile unit: <name of the last one>.