Master the fundamental concepts of x86 assembly (intel syntax) through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksSingle Instruction Multiple Data (SIMD) is where ffmpeg, BLAS, and simdjson get their speed. SSE uses 128-bit xmm registers: four float lanes in parallel. AVX widens to 256-bit ymm; your exercise stays with SSE basics in NASM.
nasmLoading…
align 16 in .dataA scalar loop issues one add per element; SIMD issues one add for four. Auto-vectorization in Clang often fails on pointer aliasing, which is why hot paths stay hand-written.
For this exercise, you will add two four-element float arrays with SSE and store the result. This task asks you to verify alignment before movaps, because unaligned SSE on older cores faults and on modern cores still hurts throughput.
Keep the relevant man page, ABI doc, or Rust reference chapter open while you work. When your output disagrees with the reference implementation on the same machine, the mismatch is usually an alignment rule, an off-by-one terminator, or a register slot you misread in GDB. Skim the official documentation for the tool or ABI named in the exercise; the prose changes, but register roles, syscall numbers, and ownership rules stay stable across releases.
Add four 32-bit integers to four others in one SSE2 instruction. Write in x86-64 NASM:
nasmLoading…
paddd, and store the result to out.movdqu for the loads and the store: the arrays are not guaranteed to be 16-byte aligned, and movdqa faults on an unaligned address.int: 2147483647 + 1 is -2147483648.The harness reads both arrays, calls vec_add4, and prints out.
Eight integers: a[0..3], then b[0..3].
The four sums separated by spaces: 11 22 33 44 for 1 2 3 4 10 20 30 40.