Master the fundamental concepts of instruction set architecture through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksx86 instructions are byte sequences built from prefixes, opcode, ModR/M, SIB, displacement, and immediate fields. Disassemblers walk these bytes; encoders emit them. Understanding the format is prerequisite to writing tools like yours.
cLoading…
The ModR/M byte names two registers and an addressing mode:
For example, 89 C3 is MOV EBX, EAX in 32-bit mode: opcode 89, ModR/M places source and destination registers.
For this exercise, you will hand-encode three instructions and verify with objdump -d or the mini disassembler. This task asks you to treat bytes as structured data, the same skill malware analysts use when carving shellcode.
Keep the relevant datasheet, ISA manual, or architecture textbook chapter open while you implement. When your output disagrees with the reference trace on the same program, the bug is usually a mis-decoded opcode, a stale register read, or a flag bit left unchanged after arithmetic.
For this exercise, you will use those habits while implementing the requirement in the starter code. Microarchitectural product names change across CPU generations, but the control ideas (fetch, bypass, cache lines, vector lanes) stay stable enough to debug from first principles.
Write a small x86-64 assembler for a handful of instructions and show how every byte is built: the REX prefix, the opcode, the ModR/M byte and the immediate. After this you'll be able to read objdump -d byte columns, and you'll see why x86 instructions are anywhere from 1 to 15 bytes long.
| code | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| 32-bit | eax | ecx | edx | ebx | esp | ebp | esi | edi |
| 64-bit | rax | rcx | rdx | rbx | rsp | rbp | rsi | rdi |
A 64-bit operation gets the prefix REX.W = 48. Both operands of a two-register instruction must be the same size.
| Instruction | Encoding |
|---|---|
OP dst, src (register, register) for add 01, or 09, and 21, sub 29, xor 31, cmp 39, mov 89 | opcode, then ModR/M = 11 src dst |
OP r, imm for add/or/and/sub/xor/cmp (/digit = 0, 1, 4, 5, 6, 7) | −128..127: 83 ModR/M(11 digit r) imm8 · else if r is eax/rax: short form 05 0d 25 2d 35 3d + imm32 · else 81 ModR/M(11 digit r) imm32 |
mov r32, imm | b8+r imm32 |
mov r64, imm (imm fits in signed 32 bits) | 48 c7 ModR/M(11 000 r) imm32 |
inc r / dec r | ff ModR/M(11 000 r) / ModR/M(11 001 r) |
push r64 / pop r64 | 50+r / 58+r |
ret / nop | c3 / 90 |
ModR/M is mod(2 bits) reg(3) rm(3). Immediates are little-endian two's complement. Immediates are decimal, possibly negative, in the signed 32-bit range.
One line per instruction, bytes in lower-case hex:
cLoading…
The bracket lists, in order, REX.W (if present), op= (as b8+r/50+r/58+r for the register-in-opcode forms), modrm=mod.reg.rm in binary (if present) and imm8=0xHH / imm32=0xHHHHHHHH (if present). The instruction is echoed as mnemonic op1, op2, normalised to single spaces. A line that can't be encoded prints LINE => error. After the last line print total: N bytes.
Input:
cLoading…
Output:
cLoading…
81).Hidden tests cover negative imm8, the 81 form for a non-accumulator register, mov with immediates in both sizes, inc/dec/pop/nop, and lines that can't be encoded (mixed sizes, push of a 32-bit register, unknown mnemonics).