Master the fundamental concepts of cpu design through this focused micro-challenge.
You have read the whole brief, and the concepts above stay free on every task. Writing and running the code needs a plan.
Three hints are available for this task, revealed one at a time inside the code workspace so you can struggle productively before seeing them.
Every task includes starter code, theory, and hidden tests so you can implement and verify locally in the browser.
How it worksForwarding (bypassing) routes a computed result from EX, MEM, or WB directly to the EX stage input muxes, avoiding a register-file round trip. This kills most data hazards without stalling.
cLoading…
For example, after ADD R1, R2, R3 immediately followed by SUB R4, R1, R5, the SUB in ID needs R1 while ADD is still in EX. Forwarding picks the ALU output from EX/MEM instead of reading the old register value.
For this exercise, you will implement muxes that select the youngest matching producer. This task asks you to compare IPC with and without forwarding on synthetic hazard-heavy code.
Keep the relevant datasheet, ISA manual, or architecture textbook chapter open while you implement. When your output disagrees with the reference trace on the same program, the bug is usually a mis-decoded opcode, a stale register read, or a flag bit left unchanged after arithmetic.
For this exercise, you will use those habits while implementing the requirement in the starter code. Microarchitectural product names change across CPU generations, but the control ideas (fetch, bypass, cache lines, vector lanes) stay stable enough to debug from first principles.
Add forwarding (bypassing) to the 5-stage pipeline. Instead of waiting in ID until a result has been written back, an instruction in EX takes the value straight from the pipeline latch that holds it. That removes almost every data-hazard stall. The exception is load-use: a loaded value only exists after MEM, so the instruction right behind a load still stalls one cycle. Run the same program with forwarding on and off and compare.
The first line is forwarding on or forwarding off, followed by a program for the ISA from the previous task:
| Instruction | Reads | Writes |
|---|---|---|
LI Rd, imm | - | Rd |
ADD Rd, Rs, Rt / SUB Rd, Rs, Rt | Rs, Rt | Rd |
LW Rd, off(Rs) | Rs | Rd |
SW Rt, off(Rs) | Rt, Rs | memory |
BNEZ Rs, label | Rs | - (decided in EX) |
HALT | - | - |
Registers R0: R7 and 16 memory words start at 0; addresses are (Rs + off) mod 16. Instructions are named I0, I1, … by position.
Same as the previous task: IF, ID, EX, MEM, WB. Registers are written in the first half of a cycle and read in the second. A taken BNEZ in EX flushes ID and IF, and this beats a stall. Fetching stops after a HALT is fetched and restarts on a taken branch. The run ends after HALT leaves WB, or when the pipeline is empty. The limit is 300 cycles.
LW whose destination ID reads (load-use). For the instruction in EX, each distinct source register (in operand order: Rs, Rt; Rt, Rs for SW) is forwarded from the youngest producer:
fwd Rn<-MEMfwd Rn<-WBcLoading…
(-- is empty or a bubble; the forwarding events come first.) Then:
cLoading…
On the cycle limit, print cycle limit reached instead of the first two summary lines.
Input:
cLoading…
Output:
cLoading…
With forwarding off the same program takes 12 cycles with 4 stalls.
Hidden tests cover a load immediately followed by a use (one stall even with forwarding), a store whose data register is forwarded, and a countdown loop whose branch uses a forwarded value.